04 · Evidence and roadmap

Capability grows in a logical sequence.

This is not a calendar or a collection of feature promises. Each stage creates the evidence, control, or operating capacity required before the next stage should begin.

01
Verified foundationWorking now

Durable knowledge, governance, validation, and recovery

The organization needs an enduring memory and authority layer before orchestration can safely become more capable.

18interoperability tests passed41 / 41recovery items verified1,700+source analyses routed in the July 2026 baseline

*Evidence basis: self-audited project records. The 18-test suite covered provider-neutral boot, memory routing, replay, conflict, rollback, and source-class controls, with zero failures or skips. Recovery verification matched all 41 manifested files with zero missing, mismatched, or extra items. The July 15, 2026 baseline counted 1,769 source-analysis notes. These figures demonstrate the working foundation, not an enterprise deployment or measured business return.

Capability
Provider-neutral context, explicit information authority, bounded workflows, human authority, trace, replay, and recovery.
Next-stage condition
The foundation must remain stable while multiple routes use the same rules, evidence, and decision history.
02
In developmentUnified orchestration

One governed path from objective to eligible capability

Users should not have to make architecture, privacy, context, and model-selection decisions one prompt at a time. Ludwig is the planned orchestration function that presents one governed interaction plane while coordinating the capabilities behind it.

Dedicated communications channelOne place to state objectives and receive decisions
LudwigPolicy, context, routing, validation, and trace
Controlled runtime environmentMultiple specialist agents, models, and tools
Capability
State the objective through a dedicated communications channel, apply policy, assemble minimum authorized context, coordinate multiple agents inside a controlled runtime, and select the simplest eligible route.
Requires
The durable foundation, authority model, and evidence qualifiers established in Stage 01.
Next-stage condition
Routing decisions must be inspectable and reproducible before cost or quality comparisons become meaningful.
03
Planned evaluationMeasured choice

Reproducible cost, quality, latency, and rework baselines

The expensive mistake is not paying for frontier intelligence. It is paying for frontier intelligence when the task never required it.

Capability
Compare eligible routes using tokens per accepted result, cost, latency, escalation, quality, rework, context size, and verification need.
Requires
Stable, recorded orchestration decisions from Stage 02.
Next-stage condition
Claims about savings or better routing wait for a baseline, repeatable method, and accepted results.
04
ExploringSpecialist capability

Local and specialist routes for bounded work

A narrow task may be handled more privately or reliably by a specialist route, while difficult synthesis and consequential uncertainty remain eligible for stronger review.

Capability
Task-specific local adapters, deterministic tools, specialist agents, and selective frontier adjudication.
Requires
Evidence from Stage 03 showing where specialization is justified.
Next-stage condition
Specialist routes must meet quality, privacy, traceability, and recovery thresholds.
05
Longer-term directionBounded improvement

Research and procedural learning with human approval

Observed friction should be able to generate research and improvement proposals without silently granting a system permission to rewrite itself.

Capability
Move from observed friction through research, proposal, controlled testing, human approval, monitored release, and recovery.
Requires
All earlier stages, plus explicit evaluation and rollback boundaries.
Boundary
This automated loop is roadmapped and inactive. Current governance and review patterns do not imply autonomous self-modification.