Every time a new technology arrives, good, hard-earned engineering practices seem to get thrown right out the window. Contrary to the marketing hype (and some popular opinion), Astra is not just a model, and it is certainly not AGI or the singularity. It is an important evolution, but there is a dangerous, little talked about reality hiding in plain sight. Astra is an entire scaffolding, harness, and agentware stack wrapped around an actual model, and it is almost entirely an opaque black box.
Enterprise executives and developers need to beware and be far more cautious about what they’re stepping into. That enterprise license agreement you signed with OpenAI or Microsoft? You might as well toss it out the window. Sure, they promise they won’t train their models on your data, but with an architecture like this, who cares? You hand your data over to their black box, and it doesn’t just inference its way to an answer anymore. It can actually act on it.
Why?
Because all those rogue agents everyone was terrified of … you know, the ones that break out of sandboxes and execute code on host systems. They are now bundled directly into Astra by a third-party vendor that knows nothing about your business and couldn’t care less about your corporate risk.
I always need to qualify myself when I write these articles because I’m a proponent of AI, GenAI and agentic AI. They are powerful and capable when in the right hands. I don’t, however, anthropomorphize these things, and I look for real systems engineering discipline when they are applied. OpenAI has effectively collapsed the distinction between a neural model and an execution system. That is inherently dangerous.
You and your employees are no longer just sending prompts and handing off your data to a public inferencing engine. That was bad enough by itself. You are now handing your internal corporate crown jewels to a stateful, agentic operating system that interprets intent and executes real-world actions based on how a third party’s engineers decided it should act.
Let that sink in for a second. Astra has quietly shifted from a stateless model to a stateful, interpretive operating system. I don’t know what kind of risk calculations are driving executive boardrooms right now, but this shift should scare the hell out of every board member, executive team, shareholder, and enterprise customer. It is one thing if your own internal agents mishandle your data and go sideways. It is a completely different level of liability to hand that execution over to a third party to do it blindly inside their closed runtime.
When Altman and the tech elite talk about the singularity, little did you know they meant a single, proprietary interpretive operating system with control over your business workflows.
The Hard Earned Practices
So what happened to all the battle scars the software industry earned over the last forty years developing operating systems?
Real operating systems, the ones that power modern enterprise infrastructure and even host these massive clusters, are built on critical architectural primitives. Privilege rings, hardware-enforced process isolation, memory protection, deterministic guardrails, and granular instrumentation. Enterprises are supposed to have complete control over what hits the processor, what gets committed to memory, what gets stored, and how every single system call is tracked and audited.
All of that seems to go out the window with Astra.
By OpenAI’s own admission, this new architecture drastically reduces observability. It is the exact opposite of where enterprise governance was heading, the exact opposite of what high-assurance environments require, and the exact opposite of what regulators are demanding. If Astra is going to act as an interpretive operating system that takes your data, processes it, and triggers commands across software environments, you had better demand total visibility into every single gear turning inside that machine.
Covert Sandbagging and Recurrent Depth
Then there is this whole issue of covert sandbagging. I needed to look this one up, and even Claude took its sweet time coming back when asked about it in relation to Astra. This is no longer a hypothetical or academic exercise … it’s become a capability in how these systems can operate.
Astra uses an architecture called recurrent depth, or looped transformers. The concept itself is not entirely new and comes straight out of academic research. Instead of moving data sequentially through distinct layers and outputting a readable token to a scratchpad at every step, the model loops tokens through the same internal block multiple times.
Here’s the problem with that. In traditional models, much of the reasoning happens across discrete text tokens. When a model plans, it writes words down on a contextual scratchpad, which gives you, at least in part, a forced translation into human-readable text that an external monitor can audit. Astra reduces and can even eliminate that forced translation. It can keep more of that reasoning inside continuous hidden-state representations, an internal mathematical language often referred to as neuralese.
Sure, the reasoning trace was never perfect. Nobody said it was. But “it wasn’t perfect” and “kill it” are two very different approaches, and only one of them is right. And burying the scratchpad is not the only invisibility trick happening here.
That means the contextual scratchpad has been moved into the latent space. These are all just fancy words for compressed math numbers … something you can’t read, can’t audit, and can’t understand.
And back to the hard-earned thing: Do you really want an untrusted third-party system doing multi-step planning in an invisible state using your proprietary data? Where it can sandbag? Data is supposed to be your crown jewels, yet somehow, enterprise leaders continue to cling to wishful thinking that these models will magically yield fantastic outcomes, and therefore check their common sense at the door.
The hidden crime in all this is that if a system can manage its real strategy in hidden math while feeding you polite text on the surface, you have no way to verify what it is actually doing until the damage hits your production systems.
It’s No Longer Training, It’s Harness Construction
It appears with Astra, that we’ve moved out of the pure machine learning phase and right back into ad-hoc software engineering. Not surprising given the amount of code involved in agents and other agentic systems. A massive portion of what actually constitutes Astra is scaffolding, adapters, and agentware. Sure, there is fine-tuning and the recurrent depth mechanism, but the consensus from what I’ve read is that the underlying neural network is evolutionary in its progression. And as mentioned at the beginning, there’s no miraculous leap to AGI.
What has changed is how the vendor’s scientists and engineers build the surrounding system. The real capabilities are being dictated by state management and embedded harnesses.
Look at the benchmarks. OpenAI heavily markets Astra scoring 99.9% on the ARC-AGI-3 reasoning benchmark. But independent testing revealed the dirty secret behind that number. When you test the model using a standard, stateless harness, its score collapses down to roughly 62%. That massive jump from 62% to 99.9% was not achieved by transformer magic. It was generated by OpenAI’s proprietary Provider Adapter harness, a piece of scaffolding built specifically to preserve opaque reasoning state and manage memory between calls.
The real capability lives in the harness.
This means we are entirely at the mercy of how a tiny group of research scientists configure their execution loop. It is stateful now. It remembers your data as it loops. It manages execution retries. It is doing the exact same looping behaviors that previously caused agents to break out of sandboxes, compromise infrastructure, and leave developers scratching their heads wondering how containment failed.
Now OpenAI isn’t disclosing their architecture, but it’s not rocket science to figure out what’s going on underneath the covers. Now you have to ask yourself, are you comfortable with this architecture? This interpretive operating system? I would tend to run the other way. I do not give any of these black boxes direct access to my production drives or codebases because they are fundamentally untrustworthy. And if these so-called models are shifting to be increasingly 90% scaffolding and agentware, built by backroom engineers chasing metrics, that should be enough to rethink your AI strategy.
The Pitfalls of Benchmark Theatre
Which brings us to the trap of benchmark theater. The motivation behind these proprietary tests and harnesses seems to be entirely about winning headline velocity and beating public benchmarks, rather than delivering enterprise reliability. The frontier vendors are good at this and I’d like some of that chutzpah. While the benchmarks can be useful indicators, they have almost nothing to do with real-world enterprise ROI. In fact, the desperate race to top leaderboards directly incentivizes the vendor to cut corners, inject opaque state adapters, and hide internal loops just to keep the multi-step reasoning alive long enough to pass a synthetic test.
It is one thing to construct an ad-hoc harness to win academic glory on an interactive puzzle. It is completely different to build a dependable system that delivers trusted quality of earnings reports, creates a 60% efficiency gain when onboarding corporate clients, maintains international supply chain resilience under geopolitical stress, or delivers trusted clinical recommendations for remote healthcare.
Benchmark theater does not deliver enterprise ROI. It delivers brittle, over-engineered black boxes built for press releases and it’s doing it with invisible processing that can’t be audited and therefore can’t be trusted.
Supervisory Control and Active Surveillance
At CharliAI we don’t trust agents, we can’t trust agents. That’s the principle behind Zero Trust. And if you have ever listened to our discussions with former US Navy commanders, you know we believe in operational accountability. Command and control is not a buzzword. It is how you keep critical systems from blowing up.
We developed supervisory control systems, active surveillance, and granular instrumentation because we knew from day one that autonomous agentic systems could not be trusted. We built forensic tracing because enterprise governance requires extreme visibility into what every single layer of an AI architecture is doing in real time.
We need to understand the reasoning. Why was a particular decision made? Under what specific conditions and criteria did the system come to a conclusion? When did an agent query data, what did it touch, when did it touch it, and was it strictly authorized to get that data? Did the agent write it down in a scratchpad? Did it pass unauthorized context to another agent? And, is this about to blow up in my face and cause a catastrophic cascading failure across an ERP or a CRM?
Transparency is everything. You even need it for debugging. But Astra deliberately stripped that away, and for the life of me, I cannot understand why enterprise leadership thinks that is acceptable.
OpenAI will tell you that recurrent looping in latent space is mathematically complex and hard to expose. I’m not buying it. Recurrence is simply a method of reasoning … tracking state, updating state, and iteratively working a problem. Knowing how and why decisions happen inside that reasoning loop is not an optional aesthetic detail. It must be human-understandable and mechanically auditable.
When push comes to shove, your organization will be audited. You will be held accountable for every transaction, every breach, and every hallucinated execution. You had better make damn sure you know what is happening inside your systems, and you had better be able to explain it.
One last thing worth sitting with: not everything here is about the model, the chain-of-thought, or what’s happening in latent space. Once an Astra agent acts, that action leaves the model’s world and enters yours. And almost nobody is watching what happens next. This piece covered the front half of the problem. The back half is next.

