Validation and assurance · runtests_ts

Proving the controls, not describing them

The infrastructure definition creates the environment. The validation system verifies it. Keeping those separate allows failures to be isolated rather than mixed together.

The validation question

DRE should progressively replace assertions with observation. “The architecture is private” is a claim. “Public network access is disabled, the private endpoint exists, private DNS resolves correctly, an authorized workload can invoke the approved model, an unauthorized path cannot, and the result is captured in a repeatable test record” is evidence.

The validation sequence

  1. Static validation — deployment location resolves, naming follows the legalfirewall convention, required resources are represented, network definitions are valid, declarations compile.
  2. Controlled test deployment — deploy into an authorized, attributable test resource group.
  3. Configuration assertions — inspect deployed properties and compare them with the expected contract.
  4. Connectivity testing — prove the intended paths work.
  5. Data and model integration testing — exercise Cosmos DB and the approved model deployment.
  6. Evidence capture — retain a reviewable record of the run.
  7. Cleanup and repeatability — remove test resources and verify nothing was left behind.

Infrastructure contract assertions

ResourceAsserted properties
Virtual networkAddress space is 10.0.0.0/16; AppSubnet exists; AppSubnet uses 10.0.1.0/24.
Key VaultVault exists; Standard SKU is provisioned; deployment flags match the declared infrastructure.
Cosmos DBAccount exists; type is GlobalDocumentDB; Standard offer configured; primary location matches the deployment region.
Azure OpenAIAccount exists; S0 SKU; custom subdomain established; publicNetworkAccess equals Disabled.

Negative validation matters most

It is not enough to prove that an approved application can reach Azure OpenAI or Cosmos DB. The completed environment should also prove that the paths which are supposed to be prohibited actually fail — and that the failure is observable. That turns architectural intent into testable security behaviour.

Testing the two processing layers separately

The deterministic engine is especially suitable for repeatable coverage because its rules and expected outcomes are explicit: clear positive and negative examples, multiple triggering conditions, boundary conditions, multiple violations in one narrative, and rule-version behaviour.

Consensus validation requires a different method because the layer is probabilistic. It splits into two testable questions: what did each agent return, and given those votes, did deterministic routing produce the expected system action? Keeping those separate makes failures diagnosable.

Evidence and repeatability

Each meaningful run should produce a reviewable record: test ID, environment, infrastructure version, application version, rule-set version, model configuration, input dataset version, individual outcomes, expected denials, failures, and cleanup result. A successful deployment that leaves uncontrolled resources, credentials, storage, or cost behind is not a completely successful validation cycle.

Production readiness is an evidence state

Not a calendar milestone. The current system should not be described as having met production criteria until the corresponding implementation and evidence exist.