The Frontier Freeze: How AT&T Grows Its AI Usage on a Flat AI Budget

AT&T just did the thing every CFO has been wanting someone to do first. It told the frontier AI labs that their line in its budget is staying flat for the next several years, and then it kept growing its AI usage anyway. About forty percent of the AI queries from its hundred thousand employees already run on cheaper open-weight models working behind the scenes, the target is sixty to seventy percent, and the checks to OpenAI and Anthropic don't get bigger while all of it climbs. We call the play the frontier freeze, and the reason it should have your attention is simple. One of the largest corporate AI deployments in America just proved you can stop writing bigger checks without slowing anything down. ‍

What did AT&T actually do?

It put a router between its employees and the model bill. ‍

The mechanics are less exotic than they sound. When an AT&T employee asks the company's AI assistant something, a routing layer decides which model answers. Routine work lands on open models like Meta's Llama and Nvidia's Nemotron that cost a fraction per token. Hard problems escalate to the frontier models the company already pays for. The employee never sees any of it, they just get an answer. On some coding workloads, AT&T reports the routing cut costs by more than half with a quality drop in the low single digits, the kind of trade that only exists because someone sorted the work first. The models trail the frontier by six to ten months, and for most everyday tasks, six to ten months behind the best is still far more capability than the task needs. That last idea is one we've argued since the frontier tax, and AT&T is now the largest public proof of it. ‍

What is the frontier freeze?

It's a budget posture. Cap what you pay the frontier labs at today's level, and let all growth land on cheaper capacity underneath.

Think about how you already staff work. Your most expensive people don't review every email, they get pulled in when the stakes earn their rate, and the rest of the work runs on the team. The freeze applies the same logic to models. Frontier models keep doing the work that genuinely needs them, at roughly the spend you've already accepted. Drafting, summarizing, classifying, extracting, and the long tail of routine steps run on the cheaper tier. Nothing about usage gets frozen. Workflows grow, experiments multiply, agents get added, seats expand. The only thing that holds still is the price tier all that growth defaults to, and the budget conversation changes shape completely. Instead of explaining why the AI line doubled again, you're showing capacity growth on a flat premium spend, which is exactly the deal that gets AI programs renewed instead of clawed back. ‍

Do you need your own servers to run this?

No, and this is the part most leaders get wrong about open models.

Open weight doesn't mean self-hosted. The same open models AT&T uses rent by the token through an API, from ordinary cloud providers, exactly the way you rent a frontier model today. That means the freeze starts as a pricing decision, and the work is a workflow change, sending each task to the cheapest model that handles it well, before it is ever an infrastructure project. The realistic path runs in stages. First, swap your routine, low-stakes workloads onto rented open models and bank the difference, with nothing new to operate. Second, add a routing layer you control so the sorting happens automatically instead of by policy memo. Third, and only when volume or privacy justifies it, host the heaviest lanes on your own infrastructure the way AT&T runs some of its models in its own data centers, a decision that lives on the weight line and deserves its own analysis. Most companies can capture the bulk of the savings at the first two stages and never build a server room.

Isn't this the budget cap we warned against?

No, and the difference is the entire point. A cap rations intelligence. A freeze refuses to overpay for it. ‍

We've written about the ambition cap. When you cap total AI spend, your team stops trying things, and the cost of the experiments that never happen dwarfs the token savings. The freeze caps something different. Total consumption grows and none of it hits a ceiling, because new volume defaults to the cheap tier and escalates only when the task earns it. Under a cap, an employee with a new idea hits a wall. Under a freeze, they hit a router. Same discipline on the ledger, opposite effect on the floor.

Where does the freeze go wrong? ‍

When the sorting gets skipped, because then it really is just a cap with better branding.

The freeze has one honest prerequisite, and it's judgment rather than technology. Someone has to sort the work by what a wrong answer costs. A two percent quality drop is a bargain on internal code summaries and a catastrophe on customer-facing output, so the routing has to know which is which, high-consequence work has to stay on the strongest models, and your best model arguably belongs on review duty anyway, the argument we made in the judgment tier. Skip the sorting and you get cost control by quality erosion, and nobody in the building will trust the next efficiency initiative you propose. The other failure mode is timing. A company still finding its first valuable AI workflows should let spend grow while it learns. The freeze is for companies whose usage is climbing and whose premium bill is climbing with it, one for one, which describes more budgets every quarter. ‍

Is a frontier freeze right for your company? ‍

If your AI bill grows every time your AI usage does, yes, and you can start smaller than AT&T did.

The approach runs in order. Baseline your spend by model tier so you know what you'd be freezing. Classify your workloads by what a wrong answer costs, and move the forgiving, high-volume ones onto rented open models first. Stand up routing behind a boundary you own, the architecture move we covered in the assembly answer, so model choice becomes a config change instead of a migration. Then freeze the premium line and revisit quarterly, because the open tier keeps absorbing yesterday's frontier capability, which means the frozen dollar buys more intelligence every quarter you hold it still. AT&T proved the play works at a hundred thousand employees. It works the same way at two hundred, and the companies that build the habit now will spend the next few years growing capacity on a flat budget while their competitors explain to their boards why the bill doubled again.

‍ ‍

If you want the freeze designed for your stack, from the sorting to the routing to the boundary you own, start with an AI Blueprint or reach us at contact@theyor.com.

Previous
Previous

When Every Competitor Has the Same AI, What Actually Separates You?

Next
Next

When Your Vendor Dies, Who Inherits Your Data?