What does Dr. Jonah Tebaa: The AI Model You Approved May Already Be Gone mean in practice?
According to Dr. Jonah Tebaa, an approved AI model may already be gone because managed API vendors push silent model updates without retriggering governance sign-offs under standard material change clauses. To address this risk, Dr. Tebaa introduces a six-requirement contract checklist and a three-tier response framework. His approach mandates logging specific checkpoint identifiers, securing thirty-to-sixty-day advance change notices, retaining full input-output test pairs, and monitoring a one-to-two percent live output sample monthly to detect drift immediately.
The approval that quietly stops applying
Dr. Jonah Tebaa spends a lot of his advisory time in a specific, unglamorous room: the one where a board or a risk committee signs off on an AI system after months of testing, satisfied that the work is done. His argument, developed across a series of recent pieces on AI governance, is that this moment of relief is exactly where the real exposure begins — because the thing that got tested and the thing that keeps running are not, in his framing, guaranteed to stay the same object for very long.
To make the point concrete, Dr. Tebaa describes a composite scenario drawn from patterns across several regulated-sector engagements rather than any single named client. A bank approves a credit-decisioning assistant in the spring, after weeks of testing against a fixed set of loan applications. Months later, an analyst notices that the approval rate for one applicant segment has moved several percentage points, with no change to the bank's own underwriting policy in between. An investigation turns up the real cause: the vendor pushed two model updates in the interim, and neither one triggered a re-approval, because the contract only required the vendor to disclose "material changes to functionality" — a standard the vendor, reasonably from its own commercial position, did not consider an accuracy improvement to meet.
In his work, Dr. Tebaa treats this as the central design flaw in how most enterprises govern AI today. Governance approves a specific, tested artifact. It does not, in almost any framework he has reviewed, treat a vendor-side model change as an event that should retrigger that approval. The result, he argues, is a system that is compliant on paper and unverified in production, often for months at a stretch.
Why the gap is structural, not careless
Dr. Tebaa is careful to frame this as a design problem rather than a story about negligent institutions. He points to three forces that, in his analysis, work together to keep the drift invisible. Approval happens once, at a fixed point in time, while the underlying model, delivered as a managed API rather than a file the client holds locally, changes on the vendor's own release schedule. The contractual language enterprises rely on to catch these changes, typically some version of "material change to functionality," asks for a judgment that cannot actually be made before a change ships, since materiality is a statement about outcomes and outcomes only surface once live cases have run through the new version. And most institutions, in his observation, monitor accuracy against ground truth, a signal that can lag the underlying event by months, rather than watching the live output distribution, which would show a shift almost immediately.
None of that argues, in Dr. Tebaa's account, for institutions to stop using vendor-delivered AI, or to slow down their own operations trying to re-certify every patch. His position is narrower and more actionable: the model version itself needs to be treated as a governed object, with the same rigor an institution already applies to the decisions that object produces.
What Dr. Tebaa recommends putting in the contract
He lays out six specific requirements he believes belong in every enterprise AI vendor contract, regardless of industry. A vendor should disclose a version or checkpoint identifier at the point of approval, which the client logs as its own artifact rather than leaving buried in vendor documentation. The advance-notice clause should specify a fixed window, in his framing thirty to sixty days, triggered by any change to the model, its weights, or its fine-tuning, with no materiality test attached, since materiality cannot be assessed before the fact. Clients should negotiate the right to a pinned, frozen version for regulated workflows, and if a vendor cannot offer that, Dr. Tebaa argues the inability itself should be treated as an input to the workflow's risk rating. He also insists institutions retain the full input-output pairs from the original approval test set, not just a summary report, because a summary cannot be compared against anything later. Rolling monitoring of a sample of live output, on the order of one to two percent of daily volume in his illustrative figure, should run monthly against the approval-time distribution. And a quarterly review should exist purely to confirm that the artifact approved and the artifact in production remain provably the same thing.
Alongside the contract checklist, Dr. Tebaa proposes a three-tier response depending on how a change is discovered. A silent change found only by internal monitoring, with no vendor notice, should suspend automated decisions in that workflow immediately, pending a full re-test. A vendor-disclosed minor change, such as a patch or a small fine-tune, warrants a fast sample re-test against the original test set within roughly ten business days. A vendor-disclosed major change, such as a swap of the underlying model family, should be treated, in his words, as a new approval rather than a renewal of the old one, complete with new edge cases and formal re-sign-off.
The through-line in Dr. Tebaa's argument is that the audit trail most institutions can currently produce is only half of what a regulator or a court would actually ask for. They can usually show what was approved, and eventually what was later found running. What they typically cannot show is which version was live on any specific day in between. That gap, in his view, is not something that can be closed retroactively. It has to be built into the governance process before the question is ever asked.