Measured result, bounded claim
On eight stratified local cases, structural validity moved 0.333→0.792 and executability 0.725→0.975; recall trade-offs and lexical limitations are published.
Flagship · Applied AI systems
v1.0.0 released · reproducible evidenceA released local-first intent-to-specification compiler that treats model output as an unreliable dependency and makes provider behavior, fallback, resource use, and evaluation inspectable.
Engineering thesis
Prompt quality alone is not a systems claim. The released system combines a typed provider contract, versioned protocol, authenticated remote exposure, bounded scheduling, deterministic fallback, capability-confined file operations, and correlated outcome traces.
The three-minute reviewer path is deterministic, offline, credential-free and independent of specialized hardware while exercising the same application boundary as local and configured providers.
Released evidence
On eight stratified local cases, structural validity moved 0.333→0.792 and executability 0.725→0.975; recall trade-offs and lexical limitations are published.
Deterministic mock, supervised local inference, and OpenAI-compatible adapters pass one typed contract and error taxonomy.
53 Bun, 79 pytest, and 9 integration tests cover auth, bounds, cancellation, reconnect, supervisor failure, fallback, and path confinement.
Eleven release assets include source, SBOM, dependency inventory, checksums, provenance, and raw benchmark evidence.
Evidence boundary