Research

Akka Uses Claude to Port 65 Open-Source Projects

Akka has completed a massive experiment using Claude models to port 65 open-source projects, demonstrating how structured specifications can automate complex software modernization.

InfoQ AI1 day agoResearch
Image: InfoQ AI

Akka recently completed a large-scale test of a specification-driven AI workflow, porting 65 open-source projects to evaluate automated software modernization. The initial phase of the experiment spanned 99.3 hours and consumed 9.41 billion tokens. Out of the 65 projects, Akka reported code size or performance improvements in 57 of the completed ports. The workflow utilized Claude Sonnet and Claude Opus alongside Akka Specify to handle implementation, testing, and review.

The trial ran in two distinct phases. First, Akka analyzed all 65 repositories, generated specifications, and implemented up to 10% of each project's codebase. From there, they selected 10 projects for full implementation. The delivery harness cycled through discovery, specification, porting, benchmarking, and improvement. When comparing models, Claude Sonnet averaged 61 minutes per port, while the larger Claude Opus took 120 minutes. However, Opus was more token-efficient, consuming roughly 40% fewer tokens than Sonnet.

For software engineers, the experiment highlights a crucial trade-off between model size and constraint adherence. While Opus used fewer tokens, observers noted that smaller models like Sonnet tended to follow structured specifications strictly, whereas larger models were more prone to improvisation. Furthermore, higher effort settings increased token consumption without delivering consistent efficiency gains. Structured specifications featuring claims, evidence, and typed behavior significantly boosted first-pass success, though gaps remained regarding cross-component decisions.

Validation relied on original unit and integration tests, supplemented by automated auditors checking for security, serialization, and architectural boundaries. Performance outcomes varied wildly by project type. Applications, frameworks, and libraries generally saw improvements, but infrastructure and tooling suffered median degradation. For instance, the Dify project achieved a 143,333-times performance improvement, though the compared workloads differed, while Netflix Metaflow ran approximately 100 times slower after the port.

This is our own summary of reporting by InfoQ AI

More in Research