Every organization we speak to has an AI pilot that quietly stopped. Nobody announced a failure. It simply stopped being opened.
There is a name for it now. Pilot purgatory: the pilot is extended, demonstrated and praised, and never receives a production budget, a named owner or a date. MIT put the failure rate for generative AI pilots at 95 percent. Deloitte's 2026 report adds the part that actually stings, which is that confidence does not break on one failure. It erodes across several, and by the third stalled pilot the executives stop turning up to the reviews.
The post-mortem usually lands on one of two answers: the model was not good enough, or people did not adopt it. Both are comfortable, and neither survives much contact with the pilot itself.
Here is what a stalled pilot looks like from the inside. It asks for context you gave it last week. Two colleagues ask the same question and get two different answers, both delivered with confidence. It will summarise a contract and cannot tell you whether the decision inside it still stands. It drafts anything you like and finishes nothing that runs longer than one sitting.
None of that is a reasoning problem. A better model reasons better about the same blank page. The assistant restarts from zero every session because what it actually needs, everything your organization has already decided, committed to and learned, is not something it can reach.
Built forwards
Starts empty the day you switch it on, and becomes useful once enough has piled up. That is another quarter your pilot does not have.
Built backwards
Reads the history already in the tools you connect. The memory is populated before anyone changes a habit.
Saying an assistant needs memory is not a new observation, and we are not the only ones making it. What decides whether it works is the direction. Memory built forwards starts empty on the day you switch it on and becomes useful once enough has accumulated, which is another quarter your pilot does not have. Timer is built backwards: it connects to the tools your team already works in and reads the history already sitting in them, so the memory is populated before anyone changes a habit.
The second difference is depth. Search indexes a document so you can find it again, and stops at the file. Memory keeps going, to why the document exists, what was decided in it, who approved that, and how much weight it still carries.
Then the work starts to change shape. In August, Pallas began picking up a task that had stalled and doing the part she is permitted to do, instead of reporting that it was stuck. She prepares a partner one-pager from live research on their side and your own records on ours, and hands back a document rather than a chat message. She pulls files straight from Google Drive inside the app. When she searches your documents she names the workspaces she looked in, so an answer arrives with its sources attached.
None of that came from a smarter model. These are the same models every stalled pilot already had, working for the first time against an organization that remembers.
A pilot that stalls is not a verdict on AI. It is a verdict on what the AI was given.