· William Ortiz · 5 min read

What aerospace taught us about AI in regulated environments

  • AI
  • Security
  • Regulated
  • Process

There is a particular feeling you get the first time you watch a senior reviewer at a flight readiness gate flip past three hundred pages of evidence to land on the one paragraph that does not yet say what it needs to say. They are not being unkind. They have been doing this for thirty years, and they know the document is the thing that will outlive the meeting. If the paragraph is wrong, the system flies with a wrong paragraph attached to it. So they read carefully, and they ask the question that gets the paragraph rewritten.

That cadence, evidence in, review, rework, evidence out, was the rhythm William ran for several years on the Space Launch System program. SLS is not a domain that rewards improvisation. Every change has a number. Every number has an owner. Every owner can tell you, in writing, what changed, why it changed, who reviewed it, and what test or analysis demonstrates that the change is safe. None of that is glamorous. All of it is the reason the rocket flies.

We think about AI rollouts the same way.

Why the comparison holds

Aerospace and regulated AI are not the same problem. One launches metal into orbit. The other generates text, decisions, or actions that touch customers, employees, or the public. But the failure modes rhyme. In both, the system has emergent behavior that is difficult to fully specify in advance. In both, the consequences of getting it wrong show up later, often in ways the original designers did not anticipate. And in both, the only durable defense is a record: a chain of evidence that lets a thoughtful outsider reconstruct what you knew, when you knew it, and what you decided to do about it.

The teams we work with in regulated environments, financial services, healthcare, energy, government, are not asking whether AI is impressive. They have already seen the demos. They are asking a harder question. When something goes wrong six months from now, and a regulator or an auditor or an internal review board asks how this system came to make that decision, what will we be able to show them?

If the answer is “a screenshot of a chat interface and a vendor invoice,” the rollout is not ready.

What an evidence trail actually looks like

The first thing we install on a new AI engagement is not a model. It is the bookkeeping. That includes a few concrete artifacts.

A change-control log for prompts, retrieval indexes, and tool definitions. Treat them like code. They get versioned, reviewed, and tagged. When the behavior of the system changes, we can point to the commit that changed it and the review that approved it.

A capture of every model call that touches a regulated decision. Inputs, outputs, model version, timestamp, retrieved context, tool invocations. Stored long enough to satisfy the relevant retention requirement, and indexed well enough that a person can actually find a specific case in a reasonable amount of time. If you cannot pull up the exact transcript of decision number 84,219 within a few minutes, you do not have an audit trail. You have hope.

An evaluation suite that runs on every change and that has been agreed with the people who will eventually be answerable for the system. Not a vibes-based eyeball test. A defined set of cases, with defined expectations, that produces a defined report. The report goes into the record alongside the change.

A documented rollback. Every change to a production AI system should answer, in writing, the question of what we do if this is worse than we thought. Sometimes the answer is “revert the prompt.” Sometimes it is “fall back to the deterministic rule.” Either is fine. “We will figure it out” is not.

Change-control gates, in proportion

The instinct in regulated environments is to put every AI change through the heaviest possible review. That is not what aerospace does. SLS distinguishes between changes that affect crew safety and changes that affect the color of a label. The review burden is calibrated to the consequence.

We ask the same question on AI work. Where does this system actually make a decision that a human would otherwise be accountable for? Those are the gates that get the heavy review, the dual sign-off, the formal evaluation. The rest, the cosmetic prompt tweaks, the latency optimizations, the model version bumps that pass the eval suite, can move on a lighter cadence. Treating everything as a flight readiness review will exhaust the team and produce sloppier work on the parts that matter.

What this looks like for the people doing the work

The engineers we have worked with in regulated environments are not afraid of process. They are afraid of process that is theater. A weekly meeting where everyone signs a form that nobody reads is worse than no meeting, because it teaches the team that the controls are decorative.

The version of this we try to build does the opposite. The change log is the actual change log the engineers use. The eval suite is the suite they run on their own branch before they push. The retention store is the one the on-call engineer queries when a customer reports something odd. Compliance is not a separate track that runs alongside the work. It is the work, recorded honestly.

That is the part aerospace got right, and it is the part we are trying to bring with us. Not the paperwork. The discipline of keeping a record that someone you respect could read, six months from now, and recognize as a fair account of what you actually did.