Databricks has turned new model releases into a three-day pipeline: on launch day, all 12,000 employees get access; three days later, measured data decides whether the model is promoted or retired. Databricks described the internal process in an official blog post on September 28, 2026, aimed at two problems every company burning money on AI knows well.
The first is that "frontier" is not always real frontier. Databricks gives the example of Opus 5.0, which scored lower than the cheaper Opus 4.8 on its engineers' internal quality ratings while costing more — migrating en masse to a regressed new model hurts the company. The second is runaway cost: when Databricks gave GPT Astra to a control group with no cost mitigations, average developer AI spend jumped 60% overnight, impossible to budget for at a population of over ten thousand.
Three steps: day-one access, budget guardrails, a three-day verdict
First, make new models available to everyone on day one, but tagged "experimental." New models ship through its homegrown Unity Gateway — a central hub for governance, cost management, and observability; employees' Claude Code, Codex, and homegrown Omnigent pull the latest configuration at launch via a preinstalled UG CLI. Second, four per-user budgets backstop usage: a monthly maximum, a daily runaway limit, a "quality frontier budget" reserved for top-tier models, and an "experimental budget" for new models, preventing unproven models from being abused at scale. Third, decide after three days on three signals: private internal benchmarks (offline tasks plus head-to-head runs of old vs. new), word-of-mouth feedback from power users, and cost tracking from OpenTelemetry traces.
Measured results: Opus 5.5 down 29%, GPT-6 Sol down 48%
The week of September 21 was the live-fire test: Opus 5, GPT-6 Sol, and GPT-Luna shipped within days of each other; Databricks opened them to all employees on day one and confirmed three days later that all three sit on the cost/quality frontier, then promoted them. Comparing the same cohort's sessions against the prior week:
| Comparison | Old model (avg $/session) | New model (avg $/session) | Change |
|---|---|---|---|
| Opus 4.8 vs Opus 5.5 | $5.94 | $4.23 | -29% |
| GPT-5.6 Sol vs GPT-6 Sol | $4.52 | $2.34 | -48% |
The final call: Opus 5.5 becomes the default model for Claude Code next week; GPT-6 Sol, though cheaper, is occasionally lower quality than GPT-5.6 Sol, so it won't become the Codex default and only joins the smart router's toolkit.
The value of this playbook isn't "which model won," but a replicable enterprise answer: benchmark scores don't equal real performance on your own workloads — selection has to rest on measured data; and cost control can't rely on goodwill, it needs budget mechanics. That Opus 5.5 earned default status in Databricks' testing also corroborates the model's dual lead in cost and performance.