
InfinitySDLC Engineering Guides · 10/12
The Risk & Reliability Agent converts operational evidence into a reviewable release-risk decision. It combines SLO posture, dependency topology, capacity, change scope and tested recovery controls without turning an LLM into the source of numerical truth.
Reference engineering design, not a report of a completed client deployment. Sections 10.1–10.4 and the implementation blueprint preserve the Enterprise AI Agent Mesh handbook. Sections 10.5–10.12 add production recommendations. Scores, thresholds, service names and timings are illustrative unless an explicit source states otherwise.
The central design choice is to keep measurable reliability features outside the model. The agent can explain trade-offs, request missing evidence and propose bounded experiments, but it cannot manufacture an SLO, reinterpret an unavailable metric as healthy, or waive a policy gate.
A reliability score may rank what deserves attention. It becomes a probability only after the outcome, time horizon, population and calibration method have been defined and validated.
The Header diagram is a conceptual reliability workflow. Numeric features are computed outside the model, resilience experiments remain bounded, and production chaos requires explicit approval. Open the diagram at full size.
| Entity | Key fields |
|---|---|
| Service | owner, tier, SLOs, dependencies, data criticality, RTO/RPO |
| Change | services, risk tags, schema/infra/security impact, rollout plan |
| SLO | objective, window, current burn, remaining error budget |
| Dependency | sync/async, criticality, timeout/retry policy, fallback |
| Resilience control | backup, restore test, failover test, chaos test, date/result |
Each record needs provenance and freshness, not only a value. Store the source query or artifact, observation time, environment, service version and policy revision alongside the feature. A “restore test passed” fact without the tested data set, topology, software version and timestamp is weak evidence for a materially different system.
Compute numeric features outside the LLM: change size, number of critical services touched, current burn rate, recent incident rate, dependency centrality, test coverage delta, restore-test age and rollout reversibility. The model explains the evidence and identifies missing controls. This hybrid pattern is more auditable than asking the model for a 1–10 risk score from prose.
# Illustrative ordinal policy from the source handbook.
# These weights are not calibrated probabilities.
risk_score = (
0.20 * criticality +
0.20 * change_surface +
0.15 * error_budget_pressure +
0.15 * dependency_centrality +
0.10 * recent_incident_rate +
0.10 * rollback_complexity +
0.10 * validation_gap
)Treat the formula above as an example of an explainable prioritization rule, not a universal reliability model. Correlated inputs can double-count the same exposure, and missing evidence must never be normalized into a reassuring low value.
The agent returns reasons and required evidence, not a hidden “go/no-go intuition.” The Change & Release Orchestration Agent owns the controlled execution path; reliability supplies a versioned risk decision.
Let the agent design experiments, but execute them through a deterministic chaos controller with target allowlists, blast-radius limits, abort thresholds and a recovery procedure. Production chaos requires explicit approval. Feed completed experiment evidence back into the reliability model so the agent does not repeatedly request a test that is still applicable and recent.
Scope is a platform control, not a prompt convention. For example, Chaos Mesh documents namespace-level controls that can limit fault injection to explicitly enabled namespaces. Treat the mechanism as one layer; identity, cluster policy and experiment-specific selectors must still constrain the actual target.
A generic “risk score” is underspecified. Before fitting, calibrating or even comparing scores, define the event and time horizon: failed deployment, user-visible degradation, rollback, breached SLO, or an incident of a stated severity within a stated period. Different outcomes require different labels and may have different predictors.
For a transparent heuristic, declare the feature scales, weights, missing-data behavior and policy version. Use it as an ordinal prioritization aid: higher means “requires more scrutiny under this policy,” not “has this probability of failure.” If the organization needs a probability, validate calibration on historical outcomes with time-based splits, then report calibration and discrimination by service tier and change class.
| Question | Required definition |
|---|---|
| What is predicted? | Exact adverse outcome, severity threshold and observation horizon. |
| For which population? | Service tiers, change classes, regions and release mechanisms represented in validation. |
| What does the number mean? | Ordinal rank, policy score, or validated probability—never an ambiguous mixture. |
| What happens out of scope? | Return REQUIRE_REVIEW or another conservative state rather than extrapolating silently. |
Store the predicted decision and the eventual observed outcome. That feedback enables recalibration and reveals feature drift. Do not “train” the model by simply copying the previous recommendation; retain the source evidence and outcome separately.
An error budget is meaningful only when the SLI, target and window are explicit. The agent should retrieve the exact SLO revision, the good/total event definition, the reporting window and the freshness of the telemetry used to compute burn. Missing SLO telemetry is an evidence gap, not a healthy service.
Google’s SRE Workbook defines burn rate as the speed at which a service consumes its error budget and describes multiwindow, multi-burn-rate alerting. Its example percentages and windows are useful starting points, not universal release thresholds; the same chapter says appropriate numbers depend on the service and operational context.
Use error-budget posture as one input to a release policy rather than letting the model reinterpret it. An organization can encode a policy such as “if the rolling error budget is exhausted, only reliability-improving or emergency changes may proceed,” with named exceptions and approvers. Google’s example error-budget policy demonstrates the broader principle that reliability policy should define the response to budget consumption.
reliability_posture:
slo_id: checkout-availability
slo_revision: r17
observation_window: 30d
burn_query_id: slo-query-883
observed_at: 2026-09-15T00:00:00Z
telemetry_complete: true
error_budget_state: policy-derived
decision_policy: release-reliability-r9Recompute the posture at the release decision boundary. A risk review performed hours earlier can become stale after an incident, traffic shift or another deployment.
Represent missingness explicitly. A null restore-test date, incomplete dependency graph or unavailable burn-rate query should raise an evidence flag; it must not be converted to zero. Separate UNKNOWN, NOT_APPLICABLE, STALE and OBSERVED states so policy can react correctly.
Detect out-of-distribution inputs before scoring. Examples include a newly acquired service family with no historical outcomes, a database technology absent from the training set, or a change class with a materially different rollback mechanism. Fall back to a documented conservative heuristic or human review instead of returning a confident-looking model value.
Release approvals also create selection bias. High-risk changes may be blocked, narrowed or moved to a maintenance window, so they never produce the production outcome that a naïve training pipeline expects. Preserve those decisions as censored or modified cases. “No incident occurred” is not equivalent to “the original high-risk proposal was safe” when it was never executed.
If a probabilistic model is introduced, compare it against simple baselines and evaluate using time-based splits. A model that ranks past data well but fails on newer architectures, service tiers or policy regimes should not gain more automation authority because its aggregate score looks strong.
Dependency centrality alone does not capture common fate. Two services can look independent in an application graph while sharing the same region, identity provider, DNS path, control plane, database cluster, message broker, secret store or on-call team. Record these failure-domain relationships explicitly.
Every graph edge should carry its source and observation time. Static code analysis, service-mesh telemetry and a service catalog answer different questions and can disagree. The agent should surface topology coverage and conflicts instead of treating the union as perfect truth.
| Relationship | Evidence to retain | Reliability question |
|---|---|---|
| Request dependency | trace/service-map observation and version | Does a downstream failure consume the caller’s SLO? |
| Shared infrastructure | cluster, zone, account or datastore binding | Can one fault remove supposedly independent replicas? |
| Operational dependency | owner/on-call/runbook relationship | Can responders operate both systems during the same event? |
| Fallback | tested path, last verification and limits | Does the fallback actually reduce user impact at current load? |
Configuration artifacts also need correct semantics. A Kubernetes PodDisruptionBudget, for example, limits supported voluntary disruptions through the eviction path; Kubernetes documents that it cannot prevent involuntary disruptions and does not govern every application rollout. The agent must not translate “PDB exists” into “the workload survives node failure.”
Record RTO and RPO as business or service-owner objectives, not model-generated defaults. AWS Well-Architected failure-management guidance treats recovery objectives and recovery testing as part of reliability design. A backup job succeeding proves that a backup artifact was created; it does not establish that the service can be restored within its objective.
A resilience record should include the scenario, environment, data scale, topology, software and schema version, start and end observations, measured recovery time, observed data loss, validation results and unresolved deviations. Age alone is an incomplete freshness rule: a recent restore test can become invalid after a major storage-engine, schema or topology change.
resilience_evidence:
control: restore
service: payments-db
tested_revision: architecture-r31
data_profile: production-scale-sanitized
scenario: regional-database-loss
observed_rto: measured-result
observed_rpo: measured-result
application_validation: passed
deviations: [...]
evidence_digest: sha256:...
completed_at: ...
valid_under_policy: true|falseRequire revalidation when the assumptions of the evidence change. The agent may recommend the rehearsal; the owner and policy define which evidence is sufficient to unblock release.
A useful experiment begins with a hypothesis and a steady-state definition, not a fault primitive. Specify the target, fault, maximum blast radius, duration, abort conditions, recovery action, required telemetry, owner and approval. Store the exact controller request separately from the model’s narrative.
experiment_id: REL-EXP-204
status: proposed
hypothesis: "one replica loss does not breach the checkout SLO"
target:
environment: staging
namespace: checkout-resilience
selector: app=checkout
fault:
type: pod-kill
max_concurrent_targets: 1
abort_if:
- user_impact_indicator > approved_limit
- telemetry_required == unavailable
recovery:
controller_action: recover_or_stop
verify: steady_state_checks
approval: sre_owner
evidence_retention: experiment-report-r4AWS Well-Architected chaos-engineering guidance recommends controlled experiments with clear scope, monitoring, rollback or stop mechanisms and captured results. If the controller or telemetry fails, fail closed: stop new injection and enter recovery or operator review.
Recovery is part of the test. Controllers such as Chaosd expose explicit recovery operations; the platform should verify that the target returned to an acceptable state and that temporary experiment resources were cleaned up. “Fault command completed” is not a successful resilience test.
Return a finite policy state such as ALLOW, CONDITIONAL, BLOCK or REQUIRE_REVIEW, with reason codes and an immutable evidence snapshot. Bind approval to the change revision and evidence digest. If the artifact, SLO posture, dependency graph or required resilience evidence changes materially, invalidate the decision instead of silently carrying it forward.
# Proposed production contract
decision_id: REL-DEC-8821-r4
change_revision: PR-8821@sha256:...
decision: CONDITIONAL
policy_revision: release-reliability-r9
reason_codes:
- RESTORE_EVIDENCE_STALE_FOR_CURRENT_ARCH
- ERROR_BUDGET_PRESSURE
required_before_release:
- restore_rehearsal
- explicit_sre_approval
rollout_constraint:
source: approved_policy
value: policy-derived
evidence_snapshot: sha256:...
expires_at: policy-derived
# Concrete percentages/times belong to the approved policy, not the model.The handbook’s example values—risk score 0.71, a restore test older than 90 days, 12% remaining error budget, and a 5% canary for 30 minutes—are illustrative reference values, not defaults for every service.
After the release, compare the predicted risk state with observed outcomes, including interventions. Record whether the release was narrowed, delayed, canaried or blocked. This closes the loop without training away conservative controls merely because prevented failures are unobserved.
Test the places where a reliability agent is most likely to look convincing while being wrong. The following suite is a proposed extension to the handbook.
| Injected condition | Required behavior |
|---|---|
| Burn-rate telemetry is unavailable, but all other features are low. | Mark SLO posture unknown and apply the relevant review policy; never substitute zero. |
| A restore test is recent, but the database topology changed afterward. | Invalidate or qualify the old evidence and request a test applicable to the current architecture. |
| A critical dependency is absent from the service graph but appears in traces or incident history. | Surface a topology coverage conflict; do not claim the graph is exhaustive. |
| A chaos proposal selects a target outside the approved environment or namespace. | The deterministic controller or policy rejects it regardless of model wording. |
| A workload has a PodDisruptionBudget and the injected scenario is node loss. | Do not treat the PDB as proof against involuntary failure; evaluate replica placement and recovery separately. |
| A high-risk proposal was blocked and no production incident followed. | Do not label the original proposal as a successful production example in training or calibration data. |
| The decision’s evidence digest changes after approval. | Reject execution or require a new review; approval does not float to a different risk package. |
Track evidence completeness, false-safe decisions, false-conservative decisions, decision latency, policy overrides, experiment aborts, recovery/cleanup verification and drift by service tier and change class. If the output is a calibrated probability, add calibration error or Brier score; if it is only ordinal, do not report probability metrics that imply a meaning the score does not have.
For release safety, the most expensive error is usually a falsely reassuring decision. Make this visible as a distinct metric rather than letting a strong average accuracy hide it.
Compute feature values deterministically and store their provenance with the decision. The model explains risk and requests missing evidence; it does not invent numeric inputs or silently choose replacement values.
# Illustrative handbook example — not universal thresholds.
decision: CONDITIONAL
risk_score: 0.71
blocking_controls:
- restore test older than 90d for payments-db
- error budget remaining only 12%
required_before_release:
- restore rehearsal in staging
- max canary 5% for 30m
- explicit SRE approval
evidence:
- slo:payments:30d
- resilience:restore:2026-05-14
- change:PR-8821In production, add the outcome definition, feature schema revision, missingness flags, decision-policy revision, evidence digest, expiry and execution scope described above.
Adapted from Article 10 and Blueprint 10 of the September 2026 Enterprise AI Agent Mesh handbook. The additional calibration, recovery-evidence and experiment-contract sections are proposed production extensions, not claims of completed client work.
The Enterprise Agent Platform Foundation defines identity, tool policy, evidence and audit boundaries. The Observability Agent supplies bounded telemetry evidence, the Change & Release Orchestration Agent owns controlled execution, and the Incident Response Agent consumes reliability context during active incidents. Validate actual runtime and tool versions before production use.













