Databricks Rolls Out OpenAI Astra to 3500 Engineers

Today we gave every engineer at Databricks (N=~3500) access to OpenAI Astra. To my knowledge we are the first large company (other than OpenAI) to roll this model out wall-to-wall and I wanted to share some thoughts on why we did this and how we approach model roll out decisions in general, balancing capabilities, cost, and model choice. Before we roll out any model we use Unity Gateway to perform a robust set of offline and online (cohort based) evaluations of model quality and cost. Our findings show that Astra provides a step function improvement in developer productivity and exceeds the quality of all prior models we offer (Opus 5, Sol 5.6, GLM 5.3) especially on complex or long-horizon coding tasks. We do not see meaningful improvement on medium and low complexity coding tasks, largely because those tasks are already performed near perfectly by existing models. On the cost front, we found that providing a developer cohort (N=~200) with Astra along with no other cost guardrails resulting in a roughly 60% increase in overall token spend. For broad release, we used Unity Gateway to give developers an Astra-specific budget per month which is a subset of their overall coding budgets. These nesting budgets are designed to encourage folks to use Astra for high complexity tasks where its cost is justified but not yet as a default model for all tasks. Astra is being incorporated into our Unity Gateway Router soon, at which point developers will no longer need to actively decide when to use it vs lower cost models. We are excited about Astra and will continue to share our experience with Astra and other models.

I like the idea of giving people access to the better model without making it the default for everything. Not every coding task needs the most powerful model, and the cost difference adds up fast at this scale.

Patrick Wendell great insights. However, with the low OTPM, GLM 3.5 on max effort will blow up in an instant. As for Astra, even on low effort, it can hardly complete some simple coding tasks without blowing up OTPM. Developers need to understand the limits Databricks gives to them vs their employees 😉

That evaluation discipline is useful, especially the distinction between model quality and task suitability. For long-horizon coding, I would add operational authority to capability and cost: which repositories and environments the model may touch, which changes require review, and how quickly a bad run can be rolled back. Productivity can compound very quickly, so the controls need to compound with it.

I’ve found the hardest piece is quantifying cost per problem solved. Token cost alone doesn’t capture retries or how much compute an agent burns through on longer running tasks.

I know a few other big(ger) companies running Astra across all their employees already. It’ll be cool to see if the massive price jump of ~60% that you stated is worth it.

The nested budget approach is smart, it's basically a cost guardrail metric, solving the obvious failure mode of giving people a shiny model with no constraints. Curious what the eval actually measures on the quality side, task completion, human eval, time-to-merge? The 60% token spend jump with no guardrails is a great data point too.

What stands out to me here is less the model itself and more the operating model around it. Giving teams access to frontier AI is relatively easy; building the evaluation, cost controls, and routing needed to use it responsibly at scale is the harder problem. I also think the next important metric becomes cost per successful outcome rather than simply token spend. If a more expensive model can complete a complex task correctly in one pass versus several iterations with a cheaper model, the economics can look very different. Really interesting to see this being measured at this scale!

A strong example of treating model adoption as an engineering and business decision, not a hype cycle. Evaluating capability, cost, workload complexity, and routing together is what makes enterprise AI adoption scalable and sustainable.

How are you training ~3500 engineers to know when a problem needs Astra level capability?

Nested Astra-specific budgets inside each developer’s overall coding budget act as a practical model access control and cost governance layer steering high complexity tasks to Astra while keeping spend aligned to value. Automating this via Unity Gateway Router shifts AI governance from manual policy to routing driven enforcement. #AIGovernance #Cybernorse #FinOps #CloudSecurity #InfoSec

See more comments

To view or add a comment, sign in

Explore content categories