DoorDash Builds Internal GenAI Platform for 5,000 Users
DoorDash has scaled its internal generative AI platform to over 5,000 users, transitioning to open-weights models to slash costs and improve reliability across the company.

DoorDash has successfully scaled its internal Generative AI platform to support more than 5,000 active users, with 45 new users onboarding every day. Surprisingly, 40 percent of these users are non-engineers from departments like legal, sales, and operations. The platform began in April 2023 with a vendor-first approach using proprietary models from OpenAI, Anthropic, and Google. However, to manage high costs, quota constraints, and the one-year deprecation cycle of models like Gemini, the engineering team shifted toward hosting open-weights models in-house.
To host these open-weights models, DoorDash partnered with Modal, a Python-first GPU cloud. The company leverages open-source serving engines like vLLM and SGLang, alongside fine-tuning libraries such as Hugging Face TRL, Unsloth, and Axolotl. By transitioning from proprietary frontier models to open alternatives like Qwen3 and Qwen3 Embeddings, some internal teams achieved a 20x reduction in costs while actually increasing accuracy. Other open models utilized include GLM, Kimi, and DeepSeek. This architectural pivot has already yielded annualized savings in the single-digit millions of dollars.
A core component of the infrastructure is the LLM Gateway, which provides a unified API and SDK. This gateway allows engineers to easily swap models and configure automatic fallbacks, such as routing from OpenAI to Azure or from Claude to Amazon Bedrock. By 2025, the platform evolved to support agentic workflows, expanding into an Agent Gateway. This system manages over 25 agent projects and integrates with tools like Claude Code, Codex, and Cursor. It also hosts more than 50 Model Context Protocol (MCP) servers, connecting internal agents to Slack, GitHub, and Jira.
For AI practitioners, DoorDash's journey demonstrates that a centralized platform can dramatically accelerate development velocity while maintaining strict cost controls. The gateway architecture abstracts away complex plumbing, rate limits, and authorization policies, allowing product teams to focus entirely on building applications. By prioritizing API-first designs and robust fallback routing, developers can experiment with new models seamlessly without disrupting production services.
This is our own summary of reporting by InfoQ AI



