Business

Uber Eats slashes search latency by 50 percent

Uber redesigned its Uber Eats search pipeline to slash end-to-end latency by 50 percent, prioritizing real-world user rendering speeds over traditional backend API metrics.

InfoQ AI4 days agoBusiness
Image: InfoQ AI

Uber overhauled the search infrastructure for its Uber Eats platform, achieving a 50 percent reduction in end-to-end latency. Rather than optimizing backend API response times, the engineering team shifted its primary metric to Above-the-Fold completion, which tracks the time required to render the first screen of results with images. By implementing pagination with server-side caching and asynchronous rendering, Uber improved this above-the-fold latency by more than 200 milliseconds. An agentic coding workflow was also utilized to identify and validate further optimizations.

The performance gains came from a series of incremental optimizations across the stack. Uber reduced retrieval workloads by eliminating low-value retrieval strategies, saving about 120 milliseconds. Integrating product-level embeddings cut data lookups by more than 100 times, shaving off another 50 milliseconds. Separating ranking hydration from presentation data reduced latency by over 100 milliseconds, while removing dependencies and applying request hedging contributed savings of 35 milliseconds and 40 milliseconds, respectively. Additionally, redesigning the advertising path with column-oriented bid data, in-memory access, and reduced serialization cut latency by approximately 130 milliseconds.

These optimizations build on Uber's existing search platform, which utilizes Apache Lucene, Spark-based indexing, Kafka-based streaming updates, and Go data structure adjustments to minimize garbage collection overhead. For practitioners, this overhaul demonstrates that performance tuning is often about incremental gains. As Uber engineer Pratik Dhanave noted, the success relied on "a long list of careful decisions across the full stack." The company is now testing end-to-end microbatching, Zero Pass Ranking, and HTTP multipart streaming, with early product-based search tests already yielding a 50 percent reduction in p99 latency.

This is our own summary of reporting by InfoQ AI

More in Business