InfinitySDLC Engineering Guides · 06/12
This agent coordinates deterministic delivery systems. It should never replace CI/CD; it chooses, sequences and explains existing pipelines, approvals, feature flags and rollback mechanisms.
The production boundary is deliberately narrow: the model can assemble evidence and propose a release transition, but a separate policy and execution layer decides whether that transition is legal and performs the mutation. This keeps emergency pause and rollback available even when the model endpoint is unavailable.
Reference engineering design, not a report of a completed client deployment. Sections 6.1–6.4 and the two-phase-write blueprint preserve the source handbook. Sections 6.5–6.12 add production controls and failure semantics. The rollout percentages and thresholds shown below are illustrative examples, not universal defaults or Infinity client results.
A release approval authorizes one immutable intent. It is not a reusable permission for the agent to change the artifact, target, migration or rollout strategy after review.
The Header diagram is a conceptual release pipeline. Promotion and rollback are deterministic policy decisions around existing delivery systems. Open the diagram at full size.
Trigger the agent from a signed event such as “release candidate created,” “PR approved,” or “change window opened.” Treat the event as a pointer to work rather than authoritative state. The agent resolves the current PR, artifact, required checks, target environment, maintenance window and policy version before constructing a plan.
signed event
-> task admission
-> refresh authoritative state
-> gather release evidence
-> compute proposed release plan
-> deterministic preflight gates
-> approval bound to exact intent
-> progressive delivery
-> observe
-> advance | pause | rollback
-> independent verification
-> close release evidence packet
Persist the event identifier and source revision so duplicate delivery does not create duplicate releases. If the event and current system state disagree—for example, the referenced pull request has changed—the refreshed state wins and the stale event is recorded rather than trusted.
Keep read operations separate from proposal and commit operations. A tool named release.commit_deploy should not accept free-form parameters that let the model reinterpret an approved action. Downstream services reauthorize the task identity and target on every mutation.
DRAFT -> PREFLIGHT -> READY_FOR_APPROVAL -> DEPLOY_CANARY
-> OBSERVE_CANARY
-> {PROMOTE_25 -> OBSERVE -> PROMOTE_100 | ROLLBACK}
-> VERIFIED -> CLOSED
Any active state -> PAUSED
Any deployed state -> ROLLBACK when a hard policy gate fires
Ambiguous execution -> OUTCOME_UNKNOWN -> RECONCILE
Make the state machine deterministic. The LLM proposes a transition and explains the evidence; a policy service verifies the legal transition, required evidence and authorization. Store the prior and next state, operation ID, policy revision and release-intent digest in the transition record.
Do not collapse DISPATCHED, OUTCOME_UNKNOWN and VERIFIED into “deployed.” A workflow engine can preserve control flow, but it cannot prove what an external deployment system actually changed unless the platform reconciles the authoritative state.
Risk classification affects required evidence and autonomy; it must not quietly relax the organization’s hard constraints. A low-risk classification cannot make an unapproved production identity valid, and a high urgency label cannot bypass an approval that policy requires.
The object presented for approval should be canonical and immutable. At minimum, include the exact application artifact digest, environment, rendered configuration or manifest digest, migration package, rollout-policy revision, feature-flag plan, requested window, preconditions and expiry. Hash the canonical form and approve that digest.
release_intent:
service: payments
target: production-eu
artifact_digest: sha256:...
config_digest: sha256:...
migration_digest: sha256:...
rollout_policy: progressive-r7
feature_exposure_plan: flags-v12
required_checks: [qa-packet-184, security-gate-77]
expires_at: 2026-09-15T02:00:00Z
intent_digest: sha256:...
Use immutable artifact identities rather than approving a symbolic tag that can later point somewhere else. Kubernetes documents that image tags can move while content digests are immutable; pinning the digest gives the executor a stable artifact identity. See Kubernetes image documentation.
Approvals also need attributable identities and separation of duties where policy requires it. As one platform example, GitHub environments can require reviewers, optionally prevent self-review, and withhold environment secrets until the approval gate is satisfied. Feature availability depends on the GitHub plan and repository type; the general design rule is to enforce approval and credential release outside the model.
Immediately before commit, recheck the target, artifact digest, relevant policy revision and preconditions. If any material field differs from the approved object, invalidate the approval rather than “updating” the plan under the same approval ID.
“Error rate looks fine” is not a reproducible gate. Define the metric source, exact query or detector version, canary and baseline cohorts, observation window, minimum sample requirements, late-data policy and behavior when the signal is missing or inconclusive.
| Gate element | Required definition |
|---|---|
| Population | Which requests, users, regions or workloads belong to canary and baseline cohorts. |
| Measurement | Metric/query identity, units, aggregation and threshold owner. |
| Window | Observation duration, warm-up period and handling of delayed telemetry. |
| Missing data | Pause or fail closed for required signals; never interpret absence as success. |
| Decision | Advance, pause, abort or request human review, with evidence persisted. |
Argo Rollouts analysis provides a concrete implementation example: an AnalysisRun can succeed, fail or be inconclusive and can drive continue, abort or pause behavior. Its canary strategy can use weighted steps and traffic routing. The reference design here is product-neutral; the useful principle is that the gate has machine-evaluable semantics.
The handbook’s example below is intentionally retained as an illustration. Thresholds such as a 14.4 burn rate or a 25% latency increase are not defaults. They must come from the service’s SLO and release policy.
stages: [1, 5, 25, 100]
observe_minutes: [10, 15, 30, 30]
abort_if:
error_rate_delta > 1.0pp for 5m
p95_latency_delta > 25% for 10m
slo_burn_rate > 14.4
advance_requires: deterministic thresholds + required tests
human_approval_at: [25, 100]
# Illustrative policy only.
The dangerous failure is an acknowledgement lost after the downstream system has already mutated production. If the release executor times out after dispatch, persist the original operation ID and mark the state OUTCOME_UNKNOWN. Query the original deployment operation or the target’s authoritative state before deciding what to do next.
Every mutating operation should carry an idempotency key bound to the release-intent digest. Reusing that key with different parameters is an error. A retry is safe only when the downstream boundary provides real deduplication or the platform can otherwise prove that a second call cannot create a second business effect.
Serialize or explicitly coordinate overlapping releases to the same service and environment. A canary for revision B should not evaluate metrics while revision C is simultaneously changing the same population unless the release policy intentionally supports that composition.
Cancellation is also not rollback. Stop future actions, attempt cancellation of in-flight operations where supported, and reconcile already committed effects. A compensating action is a new controlled mutation with its own authorization and evidence.
Application rollback and data rollback are different operations. Kubernetes Deployment history restores earlier Pod-template revisions; it does not undo writes made to an external database. The Kubernetes Deployment documentation describes the workload rollback mechanism, which is why release design must model stateful changes separately.
For schema and data changes, record compatibility across old and new application versions, migration phases, authoritative writers, reconciliation, destructive steps and the point after which returning to the old binary is no longer sufficient. Prefer expand–migrate–contract patterns when the system allows them:
If an irreversible or destructive migration is required, the release packet must say so. “Rollback available” is a false statement unless the recovery design explains restore, forward repair or another verified path for the changed data.
Feature flags can decouple “the binary is deployed” from “users are exposed,” which gives teams another control surface for gradual release. But flag updates are production mutations too. Bind them to a versioned exposure plan with target cohorts, owner, expiry where appropriate, audit trail and an emergency disable path.
Keep the cohorts stable enough to interpret canary evidence. If the membership changes faster than the observation window, comparisons can become meaningless. Do not put unnecessary personally identifiable information in model prompts or flag rules; let a deterministic segmentation service resolve approved cohort definitions.
A kill switch for dangerous exposure must remain operable without an LLM. The model may recommend an earlier pause, but the hard circuit breaker is a deterministic operation accessible to the incident or release path under pre-established authorization.
A release name is not an artifact identity. The release packet should name the exact digest and, where the software supply-chain policy requires it, verify provenance against expected source and builder information before promotion.
SLSA guidance treats provenance as information about how an artifact was produced; its verification model compares provenance values against expectations. Use that concept as an independent gate rather than asking the model whether a build “looks trustworthy.”
Promote the same application artifact digest through environments when practical. Environment-specific configuration remains a separate versioned input; rebuilding “the same version” for production can invalidate evidence gathered against the staging artifact unless the build process and resulting identity are explicitly reverified.
Store the CI run, source revision, artifact digest, provenance verification result and required test packet in the release record. If any identifier changes after preflight, the dependent approval and rollout evidence must be reevaluated.
A release system is not production-ready if recovery depends on the same model service that may be unavailable during an outage. Keep deterministic runbooks and bounded executor APIs for pause, traffic shift, feature disable and rollback to a known release identity.
The recovery operation still needs verification. After rollback, verify the active artifact, routing state, health signals and any data or dependency conditions relevant to the incident. “Rollback command succeeded” is not equivalent to “service recovered.”
Some progressive-delivery controllers provide optimized rollback mechanisms. For example, Argo Rollouts documents a rollback window for eligible prior revisions. Treat such capabilities as implementation features, not assumptions; the enterprise contract is that the recovery path is defined, authorized and exercised for the selected delivery platform.
Roll-forward can be safer than rollback after a stateful or irreversible change. The release plan should identify which strategy applies at each stage instead of promising one universal undo operation.
The following is a proposed acceptance suite for the reference architecture. It tests the boundaries that a happy-path deployment demo rarely exercises.
| Injected condition | Required behavior |
|---|---|
| A mutable image tag is repointed after approval. | The executor deploys only the approved digest or rejects the mismatch; it does not silently follow the new tag. |
| The release plan changes from revision 2 to revision 3 after approval. | The intent digest no longer matches and a new approval is required. |
| A mandatory canary signal is missing. | Promotion pauses or becomes inconclusive according to policy; missing telemetry is not interpreted as success. |
| The deployment acknowledgement is lost. | The original operation is reconciled; no duplicate deployment is created merely to obtain a response. |
| A destructive data migration has completed. | The system does not claim a generic application rollback can restore the prior data state. |
| The model endpoint is unavailable during a failing canary. | Deterministic thresholds can still pause or roll back through the authorized recovery path. |
| A chat message says “approved, deploy now.” | No production authority is created unless the authenticated approval system records the required decision. |
| Target state changes after preflight. | Preconditions invalidate the stale release intent; the controller refreshes evidence rather than committing the old plan. |
Track attempted releases, blocked unsafe actions, stale-approval rejections, gate-data availability, reconciliation time, verified recovery time and independently verified outcomes by release class. Separate application-only, infrastructure and stateful-migration releases instead of averaging them into a metric that hides high-risk failures.
Change-failure or deployment-frequency metrics can inform operations, but they do not prove that this agent improved business outcomes. Report any before/after claim with a defined cohort, observation window and confounders; this guide does not claim realized deployment improvements.
proposal = release.propose_deploy(
service="payments",
version="2026.09.14-rc3",
env="prod"
)
# Returns canonical release intent, immutable action_hash,
# exact diff, policy results and expiration.
approval = approvals.request(proposal.action_hash)
release.commit_deploy(
action_hash=proposal.action_hash,
approval_id=approval.id
)
The commit endpoint accepts the approved action hash and approval reference, not arbitrary deployment parameters. It rechecks target authorization, expiry and preconditions before executing. The executor returns an operation ID that can be reconciled independently from the model session.
release_packet:
intent_digest: sha256:...
artifact_digest: sha256:...
source_revision: ...
config_digest: sha256:...
migration_digest: sha256:...
policy_revision: ...
approval_id: ...
approver_identity: ...
rollout_observation_contract: ...
ci_and_provenance_evidence: [...]
operations: [...]
final_verification: ...
Close the release only after the authoritative system state and required post-deployment checks agree with the intended outcome. A completed model conversation is never the release-of-record.
Adapted from Article 6 and Blueprint 6 of the September 2026 Enterprise AI Agent Mesh handbook. The production extensions above are reference recommendations, not claims of completed client deployments.
The Enterprise Agent Platform Foundation defines shared identity, policy, approval, audit and evaluation. The QA & Validation Agent supplies versioned quality evidence, while the Observability Agent supplies bounded release telemetry. Official implementation references are linked inline where their specific semantics matter.