
InfinitySDLC Engineering Guides · 04/12
The Environment Agent makes development and test environments reproducible, diagnosable and disposable. Its safest role is to generate and validate changes to Infrastructure as Code, then let normal CI/CD apply them.
A useful environment agent must do more than create a namespace. It must explain exactly what was provisioned, establish that the environment is usable, preserve its security boundaries and verify that temporary resources are actually removed.
Reference engineering design. Sections 4.1–4.5 and the implementation blueprint follow the Enterprise AI Agent Mesh handbook. Sections 4.6–4.11 add production recommendations supported by linked platform documentation. Examples are illustrative, not claims of completed client deployments.
An environment is not ready because a tool returned success, and it is not gone because a delete request was accepted.
The Header image shows the conceptual Git, IaC, CI/CD and environment flow. Production changes still require approved execution; “retrieve secrets” means scoped references or credentials handled by platform services, not secret values supplied to the model. Open the diagram at full size.
Give the agent read access to cloud and Kubernetes state, with write access limited to a dedicated branch or sandbox namespace. Production provisioning remains behind pull requests, policy checks and human approval.
For ephemeral environments, bounded creation rights can be appropriate when the platform independently enforces a hard lifetime, quota and network policy. The agent proposes what it needs; the control plane decides what it may receive.
The named MCP capabilities are internal tool contracts, not an assumption that all infrastructure vendors expose identical servers. Bound each response by resource scope, time window and result size.
# Illustrative request from the handbook
environment_request:
template: service-integration-test
ttl: 8h
source_commit: 4bd91e
data_profile: synthetic-v3
egress_profile: allowlisted
resources:
cpu: 6
memory_gb: 16
secrets: [ref://vault/test/payments]
The lifetime and resource values above are examples, not recommended limits for every workload. The template owner and platform policy define actual limits.
These controls must be enforced by credentials, admission and execution services. A model’s agreement to respect them is not enforcement. The production extensions below explain limitations that a checklist alone can conceal.
The handbook assigns repetitive log summarization, manifest-lint explanations and classification of known infrastructure failures to evaluated local models. Reserve a more capable coding harness for novel cross-file IaC refactors when the data policy permits its use.
Local inference can reduce data movement, but does not make logs or generated artifacts safe automatically. Apply the same controls to embeddings, telemetry, backups and debugging exports. Keep the ability to read infrastructure separate from permission to create or delete it.
A Git commit is necessary but insufficient to reproduce an environment. As a production extension, record a manifest covering the template revision, IaC runtime, provider selections, module versions, image digests, rendered configuration, data-fixture version and policy revision.
Resolve mutable tags and implicit defaults before approval. Record the actual artifacts used by the executor rather than only the friendly names requested by the agent. Preserve secret references and configuration classifications, never plaintext credentials, in the manifest.
A specific Terraform detail matters here: the dependency lock file tracks provider dependencies, not remote module versions. Pin modules separately. Otherwise two runs using the same configuration and provider lock file may still select different allowed module versions.
# Proposed manifest extension; illustrative placeholders
environment_manifest:
request_id: ENV-017
owner: team-payments
template_revision: approved-template-revision
source_commit: full-commit-digest
iac_runtime: pinned-runtime-build
provider_lock_digest: recorded-lockfile-digest
module_versions: explicit-approved-versions
image_digests: recorded-container-digests
data_fixture: synthetic-v3
policy_revision: environment-policy-r7
secret_references: [ref://vault/test/payments]
expires_at: server-calculated-expiry
inventory_record: env-inventory-017
Keep task-owned state and resources distinguishable from shared dependencies. Reproducibility is a contract about selected inputs and verification, not a promise that external services or live data will never change. Record intentional differences between environments instead of hiding them under “parity.”
A namespace is an administrative boundary, not a complete hostile-code sandbox. Kubernetes’s multi-tenancy guidance describes isolation concerns beyond namespace naming. Choose the boundary from the workload’s trust level: shared namespaces, dedicated nodes, stronger sandbox runtimes or separate clusters and accounts solve different problems.
| Boundary | Control and verification |
|---|---|
| Identity and API access | Task-scoped service accounts, bounded roles and tests that cross-namespace or cluster-wide writes are denied. |
| Workload privileges | Admission controls for privileged containers, host access and unauthorized service accounts; stronger runtime isolation where needed. |
| Network | Enforced inbound and outbound policy with both forbidden and required connectivity tested from the workload. |
| Data and artifacts | Synthetic or approved masked fixtures; scoped storage, artifact access and secret delivery. |
| Resources and lifetime | Resource quotas, task leases, external-resource inventory and verified cleanup. |
Creating a NetworkPolicy object is not proof that traffic is blocked. Kubernetes documents that enforcement depends on a supporting network implementation and that allowed traffic results from the applicable policies. Validate the effective result instead of inspecting one deny rule in isolation.
Test a forbidden destination and an approved dependency from the actual execution context. A failure caused by broken DNS is not evidence that destination controls work. Domain-based outbound requirements may need an appropriate network implementation or an egress proxy; do not assume every Kubernetes policy supports hostname allowlists.
Likewise, a ResourceQuota bounds supported resource consumption in a namespace, not the complete cloud bill. Managed databases, retained volumes and other external resources require their own inventory and spending controls.
A human approving “create a test environment” has not approved any later plan the model chooses to generate. Bind approval to the exact target account, environment, source revision, policy results and plan artifact. A material change invalidates the approval and requires a new review.
Terraform supports saving a plan for later application, but its plan documentation warns that saved plans can contain sensitive data, including values hidden in terminal output. Treat plan files as restricted artifacts with access controls and retention rules. Return a redacted summary to the agent.
Planning is not harmless parsing: it uses providers, configuration and access to infrastructure. Keep untrusted code, provider downloads, data-source behavior and runner networking within the platform’s trust and execution policy.
Terraform state locking depends on backend support. Configure it where supported and prevent competing operations. Lock contention is not a reason for the agent to disable locking or force-unlock another execution. Escalate uncertain ownership to an operator.
If apply times out, persist OUTCOME_UNKNOWN and reconcile the original execution and resources. Do not launch a second apply or an indiscriminate destroy to make the workflow look complete. Cancellation stops new work; it does not prove that earlier side effects were undone.
A temporary environment needs an owner, a server-controlled expiry and an authoritative resource inventory. Do not make cleanup depend on the same model session that requested the environment. A deterministic lifecycle service must continue after the session ends or a model endpoint fails.
# Proposed platform states, not Kubernetes built-ins
ADMITTED -> PROVISIONING -> VERIFYING -> READY
READY -> IN_USE -> EXPIRED -> DELETING
DELETING -> VERIFIED_DELETED
DELETING -> CLEANUP_BLOCKED -> operator review
Ambiguous provisioning result -> OUTCOME_UNKNOWN
Any renewal -> policy + owner + budget recheck
The Kubernetes TTL-after-finished controller cleans up finished Jobs. It is not a general expiry mechanism for arbitrary namespaces, databases or complete developer environments. The handbook’s environment lifetime therefore needs an explicit controller or equivalent lifecycle service.
Record ownership when resources are created, including external database instances, DNS records, load balancers, storage and temporary credentials. Cleanup operates only on resources the environment is authorized to own. A shared service is a dependency, not something to delete merely because a test used it.
A deletion request can remain pending while finalizers complete their work. Do not strip finalizers automatically to force a green status. Surface the blocker, resource owner and outstanding effect, then follow an approved recovery procedure.
Verify deletion against the inventory and retain a cleanup receipt. Record intentionally retained artifacts, their owners and expiry separately. Failed provisioning also enters cleanup or reconciliation; environments that never reached READY can still incur cost.
Consider an illustrative request: an integration-test environment has been created, but the payments test cannot reach its dependency. “Restart everything” is not an investigation plan. First establish which layer failed and which observations are reliable.
| Observation | Next bounded check |
|---|---|
| The workload cannot start. | Inspect scheduling and image-pull events, container exit reasons and approved configuration references. |
| The workload runs but is not ready. | Inspect readiness results and dependency health; do not treat the Running phase as application readiness. |
| A dependency name does not resolve. | Check the expected service identity, namespace and authorized DNS path. |
| Name resolution works, but connection fails. | Inspect endpoints, listening ports and effective inbound and outbound policy with allowlisted probes. |
| Infrastructure state cannot be read. | Report the missing observation and stop state-dependent conclusions; do not infer a healthy or broken resource. |
Kubernetes distinguishes Pod phase and readiness conditions. In this design, readiness additionally requires the environment’s declared smoke tests, reachable approved dependencies and the correct fixture version. A green infrastructure apply cannot substitute for these checks.
Every hypothesis should point to a timestamped observation and a discriminating check. Keep logs bounded and redact before placing them in model context. When the smallest supported fix is an IaC change, create a reviewable patch; do not mutate the running environment outside the approved path.
Persist compact investigation state so another session can resume: desired revision, observed state, completed checks, rejected hypotheses, proposed patch and unknowns. The Observability Agent can supply evidence without acquiring the Environment Agent’s write capabilities.
A provisioning demo usually exercises the happy path. The proposed acceptance suite below tests the boundaries that must remain intact when the agent, network or cloud service behaves unexpectedly.
| Injected condition | Evidence required |
|---|---|
| The agent requests a ClusterRole or another namespace. | Authorization or admission rejects the request independently of the model. |
| A forbidden destination is contacted. | The connection is denied while the required approved dependency remains reachable. |
| A secret appears in a pod log. | The test value is redacted before model context and downstream telemetry; failures block the tested path. |
| The apply acknowledgement is lost. | The original operation is reconciled without duplicate provisioning or unapproved destructive recovery. |
| The approved plan is replaced. | The executor rejects the artifact or target mismatch. |
| An environment expires with an external resource still active. | Cleanup remains incomplete until the resource is removed or its retention is explicitly approved. |
| A shared dependency appears in the deletion inventory. | Ownership checks prevent deletion and surface the inconsistency. |
| Live-state access or the cleanup controller fails. | State is reported as unknown or cleanup-blocked; admission limits prevent uncontrolled accumulation. |
Report time to verified readiness, provisioning failure rate, cleanup lag, orphaned-resource exposure, policy denials and cost over the entire environment lifetime. Include failed and expired attempts rather than counting only successful creates.
Separate model latency from queue, approval, provisioning and test time. This shows whether the bottleneck is reasoning, capacity, a dependency or an operational decision. A cost-per-ready-environment measure should include unsuccessful attempts and retained resources for a defined cohort and observation period; it is not an ROI claim.
Keep reproducible failure fixtures and execution receipts with each tested platform version. The result of an acceptance test is evidence about that version and scenario, not a universal guarantee that all possible secret formats, workloads or failure combinations are safe.
Namespace: agent-<task-id>
ServiceAccount: env-agent-<task-id>
RBAC:
reads: scoped shared QA resources
writes: sandbox namespace only
NetworkPolicy:
deny-all + DNS + registry + approved dependencies
ResourceQuota: CPU / RAM / pods / storage
TTL: expires-at=<timestamp>
Secrets:
CSI or short-lived token references
no prompt-visible secret values
The expiry notation above is a platform contract, not a Kubernetes mechanism that automatically deletes any labeled object. Enforce it using the lifecycle service described in section 4.9.
| Tool | Critical behavior |
|---|---|
env.get_desired_state | Resolve a Git SHA and rendered IaC or manifests. |
env.detect_drift | Return a deterministic plan or diff. |
k8s.diagnose_workload | Return bounded status, events, restarts, probes and log samples. |
network.run_probe | Run allowlisted DNS, TCP or HTTP probes. |
env.propose_patch | Write the agent branch only. |
ci.run_iac_validation | Execute formatting, validation, policy and unit checks. |
env.request_ephemeral | Request a template with enforced lifetime and quota. |
Avoid generic unrestricted kubectl exec. Where shell access is unavoidable, validate namespace, working directory, command class, timeout and egress, and retain the independent sandbox and credential boundary. Production resource creation flows through IaC pull requests and the normal release system.
Adapted from Article 4 and Blueprint 4 of the September 2026 Enterprise AI Agent Mesh handbook. The additional lifecycle, isolation and execution controls are proposed production extensions. Platform-specific qualifications are linked to official Terraform and Kubernetes documentation.
The Enterprise Agent Platform Foundation supplies shared identity, policy, retrieval, sandboxing and audit. The Planning & Architecture Agent provides the approved design; the QA & Validation Agent consumes the verified environment. Validate actual runtime and tool versions before deployment.













