
Akka Tests Spec-Driven AI Delivery across 65 Open Source Projects
Quick Answer
Akka's testing of a spec-driven AI delivery workflow across 65 open source projects revealed a 99.3-hour porting effort consuming 9.41 billion tokens, with performance improvements in 57 projects.
Quick Take
The study highlighted the efficiency of structured specifications and raised questions about the impact of model size and coding constraints on porting efficiency.
Key Points
- Akka's workflow involved discovery, specification, porting, benchmarking, and improvement phases.
- Sonnet averaged 61 minutes per port, while Opus took 120 minutes with 40% fewer tokens.
- Structured specifications improved first-pass implementations, but gaps remained in cross-component decisions.
- Performance varied, with applications improving and infrastructure showing median degradation.
- Validation included unit tests and checks for serialization, security, and architectural boundaries.
📖 Reader Mode
~3 min readAkka used 65 open source projects to test a spec driven workflow for AI assisted software porting, measuring specification structure, context, model and effort selection, automated validation, token consumption, and runtime performance. The initial tranche took 99.3 hours and consumed 9.41 billion tokens, with Akka reporting a lines of code or performance improvement in 57 of the 65 ports.
The experiment used two tranches. Akka analyzed all 65 projects, generating specifications and implementing up to 10% of each project's surface area, then selected 10 for complete implementation based on system characteristics and measurable results. The delivery harness cycled through discovery, specification, porting, benchmarking, and improvement. Discovery analyzed code, models, schemas, and runtime behavior, while Claude with Akka Specify handled implementation, testing, and review. A common benchmark runner compared tests, code size, and latency.

Akka delivery harness workflow(Source: Akka Blog Post)
Akka found that structured specifications with claims, evidence, and typed behavior improved first-pass implementations, while gaps in context files remained around cross-component decisions. Follow-up areas include interface enumeration, test ingestion, provenance tracking, differential testing, and adversarial testing. GitHub Spec Kit similarly structures coding agent workflows around specification, planning, tasks, implementation, and convergence. In Akka's experiment, Sonnet averaged 61 minutes per port versus 120 minutes for Opus, while Opus used about 40% fewer tokens. Higher effort settings increased consumption without consistently improving efficiency.
The finding prompted discussion among engineers following the research. Aaditya, commenting on a LinkedIn post by Tyler Jewell, CEO of Akka, wrote:
Smaller model's behavior matched modernization work he had observed, where the small model follows the spec while a larger model may improvise.
Aaditya also questioned:
Whether reductions in lines of code resulted primarily from dead code removal or from differences in the target language.
Rick Bryce, head of marketing at Avahi, raised a related point in the same discussion and suggested that constraints could influence the result. The comments add questions around whether model capability, specification constraints, or both account for differences in porting efficiency.
Cheaper model giving the tighter port is the finding worth chasing

Akka model and effort efficiency chart (Source: Akka Blog Post)
Validation used the original unit and integration tests alongside auditors checking serialization, security, error handling, PII, idempotency, and architectural boundaries. Akka added guardrails as failures exposed new issues, while reporting that additional exit conditions increased porting costs. Performance varied by project: applications, frameworks, and libraries generally improved, while infrastructure and tooling showed median degradation. Akka reported a 143,333 times improvement for Dify, but noted that the compared workloads differed, while Netflix Metaflow was approximately 100 times slower.
About the Author
Leela Kumili
Show moreShow less
— Originally published at infoq.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from InfoQ AI, ML & Data Engineering
See more →Google Cloud Workbench Notebooks Extension Connects VS Code to Google Cloud's Jupyter Notebooks
The Google Cloud Workbench Notebooks extension for VS Code allows developers to seamlessly connect their local IDE to managed Jupyter notebook environments on Google Cloud, enhancing ML workflow efficiency. This integration eliminates context switching, enabling smooth transitions from local experimentation to high-performance cloud computing.

