Modulate CEO Urges Shift Away From Monolithic AI Models
Modulate CEO Mike Pappas warns that relying on massive, general-purpose AI models for real-time tasks like voice processing is driving up enterprise costs and eroding operational trust.

The push toward massive, general-purpose artificial intelligence is creating severe inefficiencies for enterprises deploying these systems at scale. According to Mike Pappas, the CEO and co-founder of voice intelligence startup Modulate, using heavyweight monolithic models for every task introduces unnecessary computational overhead. This design flaw is particularly painful in real-time environments like contact centers, which typically process between 5,000 and 15,000 customer calls daily. Running massive models continuously across millions of annual interactions quickly becomes financially and environmentally unsustainable.
These scaling difficulties have measurable consequences for businesses. Data from a McKinsey Global Survey on AI highlights the operational toll of these inefficiencies, with 39 percent of organizations reporting increased call volumes from fraud-related inquiries. Additionally, 34 percent of companies experience longer handling times, and 29 percent report a decline in employee productivity. Because running large models continuously is too expensive, many enterprises only analyze a small fraction of their data or rely on delayed, post-incident reviews.
Beyond cost, the "black box" nature of massive models hinders real-time decision-making. In high-stakes scenarios like fraud detection, an AI system that flags an anomaly without explaining its reasoning forces human operators to choose between blind trust and costly hesitation. For voice applications, where systems must interpret continuous streams of dialogue and catch nuances like sarcasm in seconds, opaque and slow general-purpose models fail to deliver immediate, actionable insights.
To resolve these bottlenecks, Pappas advocates for a multi-stage architecture that replaces single, monolithic systems with tiered workflows. Practitioners should route routine decisions through lightweight, specialized models, reserving computationally heavy processors only for highly complex cases. This modular approach allows developers to optimize budgets, reduce energy consumption, and achieve greater transparency by tying specific decisions to distinct components rather than an uninterpretable neural network.
This is our own summary of reporting by Unite.AI



