IFCO Cuts Databricks dbt Runtime by 60 Percent
Supply-chain giant IFCO optimized its massive dbt semantic layer on Databricks, slashing daily compute costs by 58 percent and proving how large-scale data pipelines can run efficiently.

IFCO, a global circular pooling service managing hundreds of millions of reusable plastic containers, collaborated with Databricks Forward Deployed Engineering to overhaul its massive dbt-based data pipeline. By mapping dbt incremental settings directly to Delta Lake write behaviors, the team reduced the daily runtime of its core semantic layer job by over 60 percent, dropping from approximately seven hours to just two hours and 20 minutes. This optimization allowed IFCO to retire a costly nightly full refresh and cut daily compute costs by 58 percent.
The pipeline processes billions of daily tracking events to monitor asset lifecycles. To handle late-arriving data and historical reconciliation without reprocessing everything, the team implemented liquid clustering keyed to asset IDs and event dates. This enabled dynamic file pruning during merge operations, which ensures the engine only scans files containing relevant changes. For its busiest model, which consolidates tracking signals using Databricks SQL's H3 geospatial functions, the team restricted recomputations to a recent ingestion window. This change reduced the rows scanned per run by 75 percent, down from a baseline of roughly 25 billion rows, and limited the recomputed asset pool to just 3 to 5 percent of the total.
For data practitioners, the project highlights the importance of diagnosing performance issues using actual executed query plans rather than compiled SQL. IFCO discovered that its consolidation model was spilling hundreds of gigabytes to disk because of unbounded window functions and regenerated timestamps. To scale these diagnostics, the team packaged the troubleshooting steps into an automated AI-driven playbook. Operationally, IFCO transitioned from running the dbt project as a single opaque task to executing individual per-model tasks using the open-source databricks-dbt-factory library. This granular approach provides task-level visibility, targeted reruns, and automated testing on serverless compute.
This is our own summary of reporting by Databricks AI



