What if a node failure didn’t have to ripple across your cluster? As Uber’s M3DB clusters grew, shard dependencies meant that a single change could trigger recovery work across much of the cluster, making maintenance harder to run in parallel. Our engineers built a new subclustered placement algorithm to contain that impact. By grouping nodes into smaller, self-contained failure domains, we can limit shard migrations, safely run more operations in parallel, and make large clusters easier to manage. Dive into how it works. ⬇️ https://lnkd.in/gUvttaZH
About us
The Engineering team at Uber builds the technologies that power our platform and reimagines the way the world moves for the better. We thrive on the scale of our global footprint, the gratification of solving hard challenges for millions of users around the world, and being at the forefront of smart experiences and technologies.
- Website
-
https://eng.uber.com/
External link for Uber Engineering
- Industry
- Software Development
- Company size
- 10,001+ employees
- Headquarters
- San Francisco, CA
- Specialties
- software engineering, technology, programming, mobile, transportation, ridesharing, mobility, ios, android, platform, data center, infrastructure, javascript, python, golang, app, app development, and data science
Updates
-
How do you keep ML features consistent when prediction traffic reaches 8 million QPS? At Uber, logging everything would mean trillions of rows each month. So our engineers built a feature logging framework that captures the features that matter at inference and makes them available for training—while optimizing the pipeline for massive scale. The result: more consistent features, fresher training data, and a more efficient ML pipeline. See how we built it. ⬇️ https://lnkd.in/gXpxDSEX
-
9.5 million unnecessary requests. Stopped. When a service goes down, retrying a request can help. But when retries pile up across a large system, they can quickly make a bad situation worse. So, how do you retry smarter? Uber engineers built error ownership to pinpoint where a failure starts and stop unnecessary retries from spreading across our service mesh. The impact: → 9.5M unnecessary requests prevented during a major degradation → Retry storm radius reduced from up to 25 levels deep to just 3 Fewer unnecessary retries. Smaller blast radius. More resilient systems. Read more: https://lnkd.in/gqbYq4dC
-
-
📣 The first-ever H3 Conference is coming to Uber NYC. On September 23, we’re partnering with Fused to bring together the H3 community for technical deep dives, lightning talks, and conversations on how H3 is being used in production today. We’ll also hear from Uber alum Isaac Brodsky, joining a lineup of engineers and experts from across the community. If you’re building with H3, we’ll see you there! 🎟 Secure your tickets online at: https://lnkd.in/gdCjTxjd
Next week, the Uber team are partnering with Fused to host the first ever H3 Conference at Uber NYC! For those building with H3, the open source hexagonal hierarchical geospatial indexing system, this is the perfect gathering for you. Expect deep dives, lightning talks, and some great networking with engineers, researchers, and data scientists from leading organisations. We're also delighted to joined by Overture Maps Foundation and the Cloud Native Geospatial Foundation plus some incredible speakers. 🗓 When: Next week! September 23rd 2026 (from 9am to 5pm + Drinks Reception) 📍 Where: Uber, New York City Space is limited! Some come along and connect with the global community advancing the H3 ecosystem and learn how some of the projects biggest consumers are using it in production today. 🎟 Secure your tickets online at: https://lnkd.in/e2T2Awxa #H3Conference #Geospatial #DataScience #OpenSource Uber Engineering
-
-
When a request fails, how do you figure out what actually caused it? At Uber’s scale, that gets complicated fast. With thousands of services and dependencies interacting, traditional tracing can’t always give engineers the full picture quickly enough. So we built a different approach. Our engineers developed a system that automatically maps dependencies and identifies how failures move through them—helping teams find root causes faster and focus reliability efforts where they matter most. Stay tuned for more on how we’re building more resilient systems at Uber scale. Read the deep dive: https://lnkd.in/gFDWJR6n
-
-
⚡We cut Uber Eats search latency in half. How? Measure. Identify. Fix. Validate. Repeat. That's the workflow of an AI coding agent that helped us uncover and ship optimizations across the search stack. This accelerated our ability to identify bottlenecks in production, quickly test fixes, and validate their impact before shipping. It helped us turn performance optimization into a continuous feedback loop, surfacing and addressing the small inefficiencies that can compound at Uber scale. The result: faster, more seamless search—and a foundation for pushing latency even lower. Read the full engineering deep dive → https://lnkd.in/gabMgWdt
-
-
Uber Engineering reposted this
My article has finally been published! 🕺🕺 The official part: What happens when more than one orchestrator needs to control the scale of the same Kubernetes workload? Our team faced this question while rethinking how we use compute capacity during regional failovers at Uber. Adding another scaling controller sounded simple enough, but in practice we learned some hard lessons about stale caches, concurrent writers, inconsistent workload state, and safely rolling out control-plane changes. Read more about what we learned along the way: https://lnkd.in/eMHwvDdd
-
Across all agentic tools and all employees at Uber, weekly active users have grown 7x and weekly agent requests have grown 9.4x. Despite that growth, total AI spend has relatively stabilized and cost per token has decreased. In our latest blog, we break down the architecture and operational choices driving AI efficiency at scale, including better prompt caching, access to 1,000+ internal and third-party MCP servers through a single gateway, and an AI Context Graph spanning 30+ internal systems. The result? Cost per 1,000 requests for a given frontier model has fallen 34% from its peak, and cost per session is down 52%. ➡️ Learn more: https://lnkd.in/gtH6V6yK
-
-
Uber Engineering reposted this
The recording of the AI Engineer World's Fair session from Jun 30, 2026, that Adam Huda and I did on Building Blocks for Uber's Software Factory is now live! We’re moving closer to an autonomous, agent-driven Software Factory at Uber Engineering. Underneath it is an AI infrastructure stack designed to let agents safely understand our systems, take action, write and validate code, and maintain software at Uber scale. Six building blocks power the factory: * Model Gateway: PII redaction, safety guardrails, observability supporting 800+ projects and 100M+ requests/day. * MCP Gateway: Connects agents to thousands of internal APIs and SaaS tools, with token-efficient patterns like Omni MCP, CLI, and code-mode skills. * Devpods: Pre-provisioned Kubernetes environments where agents can securely build, execute, and test code. * Skills Marketplace: 3.6K skills with 30K+ executions/day, packaging reusable engineering workflows for agents. * Context Graph: 100M+ entries across 200+ node and edge types, helping agents discover context with lower token and latency overhead. * Cortana: Our AI assistant bringing everything together across Slack, CLI, and web, handling 25K+ sessions/day. We also demoed the full agentic SDLC, taking an idea from Figma → code generation → visual validation → self-healing CI → AI code review → automated code maintenance. 70%+ of Uber PRs originate from local or cloud agents, lines of code per engineer have doubled YoY, and 250+ automated migrations have touched 9M lines of code. Full video here: https://lnkd.in/gKBG9dht
Building Blocks for Uber’s Software Factory— Uday Kiran Medisetty & Adam Huda, Uber
https://www.youtube.com/
-
🔎 Looking for a few records shouldn’t mean searching through years of data. But for some export workloads at Uber, that was the challenge: small searches were turning into much bigger scans. 💡 So our engineers found a better way. By combining Apache Hudi column stats with predicate-column sorting, queries can zero in on the files that matter—and skip the ones that don’t. The result? Fewer files scanned, more focused lookups, and a more efficient way to search historical data at scale. ➡️ Read more about what we learned and where we’re going next: https://lnkd.in/g-dbJUZU
-