The Missing Piece in Your AI Strategy
For enterprise organizations investing heavily in AI, the conversation often centers on models, agents, copilots, and infrastructure. Yet one of the most important factors influencing AI success may be something much less glamorous: software performance. As AI increasingly shifts from the cloud to devices, organizations may need to revisit a discipline that has slowly faded from focus over the last decade: performance engineering.
It’s a common theme among software engineers that we feel our craft is sliding downhill. While AI may (or may not) be improving development velocity, our tradecraft has been on a downward trajectory for quite some time. Just look at Windows, or the prevalence of Electron-based applications that consume more RAM than many users realize. As computing power has skyrocketed, concerns around performance and memory usage have steadily diminished.
How We Got Here: The Performance Slide
For decades, we lived by Moore’s Law, expecting consistent leaps in hardware capability. Those improvements allowed software developers to rely on increasingly powerful systems rather than relentlessly optimize their code. In recent years, however, that trend has begun to slow.
It’s time for a change.
We need to bring performance and memory efficiency back into the common vernacular. My reasoning may be a bit of a hot take, but stay with me.
The Economics Are Changing
As the economics of cloud AI continue to evolve, rising token costs and changing billing models are forcing organizations to think differently. AI providers and their investors are no longer subsidizing our token addiction to the same degree. As a result, interest in local AI models, on-device inference, and on-premises deployments continues to grow.
And this isn’t just a cloud-side problem.
It’s bleeding into device design, operating system tuning, and now developer responsibility. If you’ve never programmed in a systems language, I encourage you to try it. We stand on the shoulders of giants, and in many ways we’ve done their legacy a disservice through our casual disregard for performance and system resources.
Big Tech Is Already Responding
Apple has leaned heavily on its M-series chips to power on-device AI capabilities. Other major technology companies are rapidly introducing their own NPU and AI inference hardware. Microsoft Build introduced a wave of AI-capable devices. But one of the largest signals of this shift, and ultimately what inspired this post, came from Apple.
I’ll be the first to admit that WWDC 2026 felt somewhat underwhelming. Yet one of Apple’s earliest announcements focused on performance improvements.
Honestly, I wasn’t expecting that.
I’m not running the latest iPhone, but I also haven’t noticed any major performance issues. The announcement sat in the back of my mind for days. Why make such a big deal about performance if everything already seems to be working fine?
Then it clicked.
AI.
The Signal From Apple That Changed My Thinking
I need at least two hands to count the number of times on-device inference was mentioned during the keynote. Nearly every AI feature or enhancement included some variation of “runs on-device” or “on-device first.”
Apple is clearly leaning into this strategy.
And you know what doesn’t work well when a device is already bogged down by bloated applications and unnecessary resource consumption?
Token generation.
When you connect the dots, a pattern emerges. Microsoft’s focus on AI PCs. Apple’s emphasis on on-device intelligence. Windows performance improvements. Specialized inference hardware. They all point toward the same reality: AI is moving closer to the edge, and performance matters again.
Microsoft and Apple cannot solve this problem alone.
This One Can’t Be Solved at the OS Level
Only so much optimization can happen at the operating system and kernel levels. Software developers need to do their part.
As we continue building AI-powered features into our applications, especially those that rely on local models and on-device inference, performance needs to become a first-class concern again. Even when a feature has nothing to do with AI, improving performance still benefits users.
Faster applications are better applications.
More efficient applications are better applications.
We’ve reached another plateau, whether you define it through performance gains, memory costs, or both. The era of endlessly throwing more hardware at software problems is becoming less sustainable.
It’s time to start doing more with less again.
How RBA Can Help
As AI capabilities increasingly move closer to the user, performance optimization is no longer just a developer concern. It is becoming a strategic business consideration. Organizations that prioritize efficient architecture, modern engineering practices, and AI-ready digital experiences will be better positioned to deliver scalable, cost-effective solutions. At RBA, we help organizations prepare their applications, platforms, and digital ecosystems for the next generation of AI-powered experiences while maintaining the performance and reliability users expect.
About the Author
Robby Sarvis
Senior Software Engineer
Robby is a full-stack developer at RBA with a deep passion for crafting mobile applications and enhancing user experiences. With a robust skill set that encompasses both front-end and back-end development, Robby is dedicated to leveraging technology to create solutions that exceed client expectations.
Residing in a small town in Texas, Robby enjoys a balanced life that includes his wife, children, and their charming dogs.