Introduction
Durable computing is a technique that makes it easier to manage resilience and reliability in distributed systems. It relies on managing state in a way that guarantees requests can run to completion, reducing the risk of system failure or downtime.
While not a new idea, it dates back to the 1970s and Jim Gray's innovations in database transactions, in an era of microservices and increasingly complex distributed systems, it's becoming foundational today. What's more, with agentic AI poised to potentially reshape the entire digital landscape, the ability to ensure systems can handle complex workflows will only grow as a concern for engineering teams.
Let's take a closer look at why this matters so much and why it's relevant in 2025.
Durable Computing: A Faster Path to Resilience
While system complexity is a key driver of durable computing, it's also fueled by growing pressure on resources, people, time, and money. Indeed, while it's entirely possible to build resilience and reliability into complex systems, doing so is expensive and requires significant upfront investment.
As Brandon Cook of Adservio explains, "organizations are turning to durable computing because they don't want to have to build all that resilience into their systems. It's a considerable effort, it takes an entire platform team to enable developers to implement these decoupled, event-based patterns."
So while a platform engineering team can provide the foundations of resilience and consistency, durable computing products and platforms enable, as Cook himself puts it,"faster delivery of independently evolving services without having to bear the cost of building much of that resilience."
The Evolving Durable Computing Tooling Ecosystem
The challenges of state management and resilience are well known, but the emergence of products that specifically address these challenges has made the term itself more prominent in industry conversations.
Born inside the tech giants
Cook considers a key factor in today's durable computing landscape to be the work of internal teams at large organizations on distributed systems challenges. There are several examples: Temporal has its origins at Uber, while Apache Airflow came out of a team at Airbnb. Open source Conductor, and its enterprise version, Orkes, were developed within Netflix. Clearly, given the scale at which these companies operate, distributed systems challenges are particularly pronounced.

Similarly, the team behind Apache Flink drew on its own experience, and the challenges faced by users, to steer the development of Restate. (The team has written specifically about the story behind that tool.)
Another durable computing platform worth mentioning is Golem. Golem was featured in the Technology Radar, Vol.31, published in the second half of 2024, which highlighted the fact that it's backed by a WebAssembly runtime. As the Golem team explains in a blog post, "WASM gives Golem Cloud the ability to make programs in any programming language indestructible, a feat that would be completely impossible with machine code, due to its highly unconstrained nature."
Managed cloud offerings
There are also durable computing offerings from the major cloud providers, with AWS Step Functions and Azure Durable Functions as key examples. However, these are more specifically focused on building and orchestrating serverless workflows within their respective ecosystems. They offer a convenient way to access durable computing if you're an Azure or AWS customer, but they don't necessarily represent the absolute cutting edge of the field.
The Limitations of Durable Computing
Durable computing offers many benefits for teams managing complex, intensive workflows. However, while there are clear advantages to bypassing the need to build resilience centrally, using platforms that offer such capabilities can create a certain form of lock-in.
The trap of opinionated platform patterns
Cook points to the risk of "getting locked into opinionated platform patterns," where the immediate benefits of a platform that seems to do everything for you from a durability standpoint end up making you inflexible.
Of course, the less opinionated the platform, the more this work has to be done by an engineering team. "I guess as you go down the stack, there are less opinionated [durable computing] frameworks. But then you have to figure out whether you want to use those platforms' cloud services, whether they're hosted elsewhere, or whether you want to host them yourself."
In other words, as always, it's a trade-off.
Idempotency and workflow interactions remain your job
"It's also not all rosy in the sense that you don't need to think about any of the key patterns," Cook cautions. "There are things like idempotency that teams still need to build into their services. There are other aspects like understanding how workflows interact with each other, so that you design your individual services or workflows or different patterns to work well within the system. There are also interesting challenges around failover and multi-region that a lot of these organizations that have built much of the platforms themselves are also starting to implement."
Being able to build durability into your agentic architecture from the outset is an important area of exploration right now.
Durable Computing in the Age of Agentic AI
One reason durable computing matters today is its convergence with agentic AI. The connection should be clear: agentic AI is a technology designed to manage complex workflows that make decisions based on diverse, scattered data sources. Their success, really, their reliability, depends on the durability of the system.
"I've seen a lot of people start to think that these durable computing frameworks could also be great for agentic architectures, because they're very workflow-oriented," Cook says. "Being able to build durability into your agentic architecture from the outset is an important area of exploration right now."
Platforms repositioning around AI
It should therefore come as no surprise that companies like Golem and Orkes are focusing on AI in their marketing. The first thing you see on Golem's homepage is the text "the AI orchestration platform," with the promise to help users "deploy secure AI applications that run reliably, retain everything in memory, and scale effortlessly." Orkes, meanwhile, describes itself as "the enterprise platform for building highly reliable applications and AI agents." This is probably where we'll see durable computing really take off, helping to mitigate some of the well-known risks of agentic AI. Just as Docker helped popularize microservices, perhaps durable computing platforms will drive the adoption of AI agents.

Putting Durability at the Heart of Innovation
Innovation in technology is largely about managing growing complexity faster and more elegantly. That challenge persists; whatever the real scale of the agentic AI revolution turns out to be, the fundamental problems of stability and reliability will always need to be addressed. In fact, these problems may well become more apparent and even more urgent for software development teams.
If that's the case, we'll likely hear a lot more about durable computing across the industry. How the field evolves is something we'll be watching closely.
Disclaimer: The statements and opinions expressed in this article are those of the author(s) and do not necessarily reflect Adservio's positions.
STAY POSTED
Get our next analyses and field notes straight to your inbox.




