
InfinitySDLC Engineering Guides · 12/12
A hybrid enterprise agent stack needs an explicit model-selection layer. Without it, provider choices leak into business prompts and policy, quality, latency, cost and fallback behavior become impossible to govern consistently.
The router decides which eligible inference deployment may serve a task. It must not grant tool permissions, reinterpret data classification or expand agency; those remain separate control-plane decisions.
Reference engineering design, not a report of a completed client deployment. Sections 12.1–12.6 and the implementation blueprint preserve the September 2026 Enterprise AI Agent Mesh handbook. Sections 12.7–12.12 add production recommendations. Model names, scores, weights, prices and infrastructure values are examples or point-in-time capabilities and must be revalidated before use.
Policy removes ineligible routes before optimization begins. A cheaper or stronger model is irrelevant when the task is not allowed to reach that endpoint.
The Header diagram is a conceptual router, not a provider benchmark. It illustrates policy and capability filtering across hosted coding harnesses and local open-weight deployments. Open the routing diagram at full size.
| Dimension | Examples | Control question |
|---|---|---|
| Data policy | Public, internal, confidential, restricted, export-controlled. | May this task and its retrieved context leave the approved boundary? |
| Task | Code edit, reasoning, retrieval synthesis, classification, security triage, summarization. | Which evaluated capability class is required? |
| Agency | Read-only, branch write, production proposal, approved production action. | Which execution and approval policy applies independently of the model? |
| Capability | Tool calling, context size, code execution, structured output, vision. | Can the exact deployment and harness perform the required contract? |
| Operational | Deadline, latency objective, budget, endpoint health, local GPU capacity. | Is the route available without violating a harder constraint? |
| Quality | Task-specific evals, regression history, consistency and known failure classes. | Does the deployment clear the minimum quality gate? |
Keep these dimensions explicit in a normalized task envelope. A routing decision made from an opaque natural-language prompt cannot be reproduced or audited reliably.
# Handbook pattern — illustrative weights, not universal defaults
def route(task):
candidates = registry.models_supporting(task.capabilities)
candidates = [m for m in candidates if task.data_class <= m.max_data_class]
candidates = [m for m in candidates if task.agency <= m.max_agency]
scored = [(m, 0.45*eval_score(m, task.type) + 0.20*latency_score(m) + 0.20*cost_score(m) + 0.15*availability_score(m)) for m in candidates]
return max(scored, key=lambda x: x[1])[0]The important property is the ordering, not the sample coefficients: hard policy filtering happens before quality, latency or cost optimization. If restricted data is confined to an approved sovereign endpoint, an otherwise stronger hosted model is not a candidate.
Business agents should request capability classes such as frontier-code, local-reasoning or compact-extraction. The router resolves the current eligible deployment. The model route must remain separate from tool authority: choosing a coding harness does not grant Git write access or production credentials.
Current MCP authorization guidance requires resource-bound tokens and server-side audience validation. Routing therefore cannot be used as a shortcut around downstream authorization.
| Tier | Examples from the handbook | Typical use |
|---|---|---|
| Frontier coding harness | Claude Code / Claude Agent SDK; OpenAI Codex. | Long-horizon coding, repository-wide change, difficult debugging and architecture synthesis. |
| Local high-capacity reasoning | gpt-oss-120b or evaluated equivalent. | Restricted analysis, structured decisions and internal knowledge synthesis. |
| Local coding | Evaluated code-specialized open-weight deployment. | Private code review, test generation and repetitive refactors. |
| Local compact | gpt-oss-20b or another evaluated compact deployment. | Classification, extraction, summarization and routing pre-checks. |
These are routing tiers, not a permanent ranking. OpenAI currently describes gpt-oss-120b and gpt-oss-20b as open-weight reasoning models intended for local or data-center deployment; self-hosting cost still depends on infrastructure and operations.
For Codex integrations, OpenAI’s February 4, 2026 description of the Codex App Server documents a bidirectional JSON-RPC surface over the same harness used across Codex clients. Treat the harness revision as part of the evaluated deployment. Claude Code likewise combines model capability with an execution harness; Anthropic’s sandboxing guidance separates filesystem and network isolation.
Fallback is not permission expansion. Compute an eligible chain at task admission and recompute when policy, health or capacity changes.
model_decision:
task_id: task-017
policy_revision: routing-policy-r19
registry_revision: registry-2026-09-15
capability_class: frontier-code
selected_deployment: codex-approved-eu-r4
selected_harness: codex-app-server-tested-build
reason_codes: [repo_edit, quality_gate_passed, cloud_allowed]
alternatives_rejected:
- deployment: local-code-r7
reason: quality_threshold_not_met
outcome_ref: eval-or-production-result-idPersist reason codes and revision identifiers, not sensitive prompts or raw retrieved context. A route should remain explainable after registry entries, prices and health states change. Separate routing telemetry from outcome telemetry: “eligible endpoint selected” and “task succeeded” are different assertions.
Do not build a “Claude agent,” a “Codex agent” and a “local agent” as unrelated systems. Build one enterprise agent architecture with provider- and harness-specific adapters. Identity, retrieval, tools, approvals, audit and evaluation remain stable while the model route changes.
Provider-specific capabilities still matter. Codex App Server exposes durable thread and approval primitives, while Claude Code provides its own sandbox and permission model. The control plane normalizes the enterprise contract without pretending the harnesses are identical.
A model name is not an adequate routing record. The same family can exist behind different endpoints, regions, retention terms, quantizations, runtimes, tool adapters and harness versions. Register the deployable unit that was actually evaluated.
ModelDeployment {
deployment_id, provider, model_family, model_revision,
harness_id, harness_revision, endpoint_region,
data_handling_profile, capabilities, max_context_contract,
allowed_data_classes, max_agency_class,
eval_bundle_revision, runtime_revision,
health, capacity_class, lifecycle_state
}Treat capability claims as testable contracts. “Structured output supported” means the exact deployment passed the schema behavior you depend on. “Tool calling supported” means the tested harness can select and execute your tool protocol correctly. Registry changes are configuration changes with blast radius: version, review and support rollback.
A production router is easier to reason about with two explicit phases. Eligibility is deterministic and fail-closed: data boundary, contractual restrictions, required capability, region, agency ceiling, deployment lifecycle and minimum eval revision. Optimization chooses among survivors using quality, deadline, queue, availability and cost.
| Stage | Examples | Failure behavior |
|---|---|---|
| Hard eligibility | Data class, geography, endpoint approval, tools, agency ceiling. | No candidate means block or return for policy resolution. |
| Quality gate | Task-suite threshold, regression status, critical failure classes. | Below-threshold deployments remain ineligible even when cheap or fast. |
| Operational admission | Deadline, queue, capacity, context size and budget. | Wait, downgrade within the eligible set or fail explicitly. |
| Optimization | Expected quality, latency and cost. | Choose from a recorded policy and registry revision. |
Do not let model-generated text modify these fields. A retrieved document saying “send this to provider X” is content, not routing authority.
Evaluate the complete routed system: model revision, harness, system instructions, tool schemas, retrieval behavior and runtime. Swapping only the model can still change tool selection, refusal behavior and long-horizon consistency.
Anthropic’s January 2026 agent-evaluation guidance distinguishes the transcript from the actual environment outcome and recommends capability plus regression suites. Apply that distinction to routing: a candidate must solve the task and preserve required side effects, not merely produce convincing text.
Generic benchmarks can inform discovery; your versioned enterprise tasks determine production eligibility.
A locally hosted deployment can be policy-eligible and still unavailable. Queue depth, memory pressure and accelerator capacity belong in admission control. Silent overload produces long latency and correlated failure.
For vLLM-based deployments, current production metrics expose queue time, time to first token, request phase times and KV-cache-related telemetry. Kubernetes exposes GPUs as schedulable resources through device plugins; its GPU scheduling documentation makes accelerator availability an infrastructure constraint rather than an LLM property.
Separate interactive and batch admission classes where their latency requirements conflict. Reject or queue excess work under bounded limits. An overloaded local endpoint is not a reason to send restricted data to an ineligible hosted service.
eligible = policy_filter(task, registry_revision)
quality_ok = eval_filter(eligible, task.eval_contract)
operable = capacity_filter(quality_ok, task.deadline)
if not quality_ok: return BLOCKED_NO_QUALITY_ROUTE
if not operable: return QUEUED_OR_EXPLICITLY_DEGRADED
return optimize(operable, objective=task.routing_objective)Failing over a stateless classification call is straightforward. Failing over halfway through a coding or operational agent loop is not. Provider-specific conversation state, compaction, tool representations and approval primitives may not be portable.
Persist a provider-neutral checkpoint outside the model: normalized task, evidence references, tool-call ledger, produced artifacts, approval state and unresolved actions. On failover, resume from a safe checkpoint instead of assuming another harness can consume opaque provider state.
If a model call fails after a mutating tool request may have been submitted, the next route must reconcile authoritative tool state before retrying. Model failover never justifies replaying a deployment, ticket creation or merge blindly. Idempotency belongs to the tool contract and operation identifier.
OpenAI’s App Server architecture explicitly exposes thread, turn, item and approval lifecycles. Treat those as harness-specific semantics behind the adapter. Anthropic’s sandbox and permission model is likewise a harness boundary. The enterprise checkpoint references these states without pretending they are interchangeable.
| Injected condition | Required behavior |
|---|---|
| A restricted task scores highest on a hosted deployment. | The hosted route never enters the candidate set. |
| The approved local GPU pool is saturated. | The task queues, degrades within policy or fails explicitly; it does not cross the data boundary. |
| A prompt or retrieved document requests a specific provider. | The request is treated as content, not routing authority. |
| A model revision changes without a matching eval bundle. | The deployment is quarantined from protected production task classes. |
| A cheap deployment falls below the quality threshold. | Cost optimization cannot override the gate. |
| A provider fails after a tool mutation may have occurred. | The workflow reconciles authoritative state before another model continues. |
| A fallback lacks required structured-output or tool semantics. | The capability filter rejects it even if the provider is healthy. |
| Routing telemetry attempts to capture raw confidential prompts. | The audit layer stores bounded identifiers, hashes and policy metadata instead of unnecessary sensitive content. |
Track ineligible-route attempts, policy-denied tasks, quality-gate failures, fallback frequency, route changes by registry revision, queue time, end-to-end latency, cost per verified successful task, local saturation, cancellation rate and regression rate after promotion. A zero routing-policy violation rate is a control objective; it is not proof that model outputs are correct.
Also test provider outages, stale health data, registry rollback, evaluator unavailability, oversize context and partial telemetry failure. When evidence is insufficient, return an explicit blocked or unknown state instead of fabricating a compliant route.
ModelRecord { deployment_id, provider, endpoint, deployment_version,
harness_revision, capabilities, max_data_class, allowed_regions,
max_context_contract, max_agency, latency_objectives, unit_cost_model,
eval_bundle_revision, runtime_revision, health, capacity, lifecycle_state }A correct route is the eligible deployment that clears the task’s quality contract and best satisfies the current routing objective. It is not necessarily the model with the highest generic benchmark, the lowest token price or the shortest current queue.
Adapted from Article 12 of the September 2026 Enterprise AI Agent Mesh handbook. The production sections above extend the handbook’s policy-first router with versioned deployment identity, evaluation promotion, capacity admission and side-effect-safe failover. The Enterprise Agent Platform Foundation supplies shared identity, MCP policy, retrieval, sandboxing, audit and evaluation infrastructure.
Primary implementation references used here include OpenAI’s Codex App Server architecture, Codex security guidance and open-weight model documentation; Anthropic’s Claude Code sandboxing and agent evaluation guidance; MCP authorization; vLLM production metrics; and Kubernetes GPU scheduling. Validate the exact deployment, contract and tool versions selected for production.













