# Cloud World Model Platform - Full Reference Cloud World Model is a multi-cloud infrastructure simulation platform that lets developers, architects, and AI agents practice and test cloud infrastructure designs without provisioning real resources - and without cloud bills. It is a Canvas Cloud AI companion product. > Simulate the cloud without the cloud. Predict behavior, analyze cost and resilience, and train reinforcement learning agents across AWS, GCP, Azure, OCI, and DigitalOcean - all in a safe, cost-free environment. --- ## Product Overview - **Platform type**: Full-stack TypeScript web application (React/Vite frontend, Express.js backend). - **Primary use cases**: - Cloud architecture education and hands-on practice (no real cloud bills). - AI/RL agent training for autoscaling and infrastructure optimization. - Chaos engineering and resilience testing. - Multi-cloud strategy evaluation and cost comparison. - Predictive capacity planning and threshold optimization. - **Simulation engine**: Capacity-aware, resource-specific performance profiles. Calculates CPU usage, error rates, throughput, connection pool pressure, and provider-specific hourly costs. - **AI explanations**: Powered by OpenAI GPT-5. Explains bottlenecks, autoscaling behavior, and architecture optimizations. Supports a beginner mode with simplified language. - **Hybrid prediction**: Combines deterministic rule-based simulation with ML-based prediction, blended via a weighted hybrid referee. --- ## Cloud Providers Supported | Provider | Compute | Database | Storage / Other | |---|---|---|---| | AWS | EC2 (t3, m5, c5, r5 families), ECS Fargate task fleets | RDS, Aurora PostgreSQL (`aurora-postgresql`), Aurora DSQL (`aurora-dsql`), Aurora Serverless v2 (`aurora-serverless`), DynamoDB, ElastiCache | S3, CloudFront, ALB | | Google Cloud (GCP) | Compute Engine (e2, n2, c2 families) | Cloud SQL, Spanner, Bigtable, Memorystore | GCS, Cloud CDN, Cloud Load Balancing | | Microsoft Azure | Virtual Machines (B, D, F series) | Azure SQL, Cosmos DB, Cache for Redis | Blob Storage, Azure CDN, Application Gateway | | Oracle Cloud (OCI) | VM.Standard.E4.Flex, VM.Standard3.Flex, Container Instances | Autonomous Database (Standard/TP) | Object Storage | | DigitalOcean | Droplets (Basic, General Purpose, CPU-Optimized, Memory-Optimized) | Managed PostgreSQL, MySQL, Redis | Spaces Object Storage, Load Balancers | ### OCI Pricing Reference Configs (June 2026, us-ashburn-1) - VM.Standard.E4.Flex: 1 OCPU + 8 GB RAM = $0.037/hr - VM.Standard3.Flex: 1 OCPU + 6 GB RAM = $0.049/hr - Container Instances (E4.Flex ~0.5 OCPU + ~6 GB): ~$0.02/hr - Autonomous AI Standard (formerly Autonomous DB): 2 ECPUs + storage blended = ~$0.789/hr ### Azure App Service Pricing Reference (East US, Linux on-demand) Rates are Azure Retail Prices API Consumption meters with a `1 Hour` unit, retrieved 2026-09-26 from https://prices.azure.com/api/retail/prices. Windows plan meters and non-hourly units are excluded: | Plan | vCPU / RAM | Hourly rate | |---|---:|---:| | Basic B1 | 1 / 1.75 GB | $0.017/hr | | Basic B2 | 2 / 3.5 GB | $0.034/hr | | Standard S1 | 1 / 1.75 GB | $0.095/hr | | Premium v3 P1v3 | 2 / 8 GB | $0.155/hr | ### AWS ECS Fargate task fleets (Linux, us-east-1, verified 2026-09-25) - Use a compute resource with `serviceFamily: "ecsFargate"` and `requestServingKind: "container"`. - Select a task allocation with an exact canonical `characteristics.size` label (for example `"2 vCPU / 4 GB"`), the equivalent `"fargate-2vcpu-4gb"` label, or matching `containerCpu` + `containerMemoryGiB` / `vcpu` + `memoryGiB`. A Fargate-prefixed label identifies the family even without `serviceFamily`. Conflicts, malformed labels, missing halves, and unsupported pairs return `INVALID_FARGATE_TASK_SIZE`; the 0.25 vCPU / 0.5 GiB default applies only with no sizing input. - Supported Linux task allocations range from 0.25 vCPU / 0.5 GiB through 16 vCPU / 120 GiB. Configure `desiredTaskCount`, `minTaskCount`, `maxTaskCount`, and `perTaskCapacityRps`; the simulator changes a bounded task fleet rather than cloning generic VMs. Valid task labels are canonicalized to the same CPU/memory fields used by runtime billing and do not produce `unrecognized_skus`. - Tasks bill while running at $0.04048 per vCPU-hour plus $0.004445 per GiB-hour on x86_64 (the default), or $0.03238 per vCPU-hour plus $0.00356 per GiB-hour on ARM64 (`cpuArchitecture: "ARM64"`), with per-second billing and a 60-second minimum for a newly started task. A service with zero running tasks costs $0. Explicit regions outside us-east-1 and unsupported architectures are rejected; neither is silently assigned a different rate. Source: https://aws.amazon.com/fargate/pricing/ (Linux us-east-1 worked example and architecture tables; verified 2026-09-25). - Startup and scale-down waits are controlled with `taskStartupSeconds` and `taskScaleDownSeconds`. At `maxTaskCount`, surplus demand is modeled as bounded queueing, latency, and errors. - Not modeled: ECS on EC2, Fargate Spot, ALB/Cloud Map/NAT/egress, image registry charges, batch jobs, service discovery, or Savings Plans. - Direct AWS RDS lifecycle keys (`extendedSupport`, `engine`, `engineVersion`, `databaseEngine`, `databaseEngineVersion`, `dbEngineVersion`) are rejected when placed on resource characteristics/top-level resource input. RDS Extended Support is not priced. The bounded SQL Server Express model below uses its own strict nested `rdsSqlServerBilling.engine` discriminator; that is not a lifecycle-support input. The simulation-level `engineVersion` response field is server-managed CWM metadata, not a database engine version. ### AWS RDS SQL Server Express T3 billing (bounded model) Use an AWS database resource with `characteristics.serviceFamily: "rds-sqlserver-express"` and a nested `characteristics.rdsSqlServerBilling` object. This contract is deliberately narrow: SQL Server Express (license included), Single-AZ, on-demand, `us-east-1`, T3 Unlimited, supported `db.t3.micro` through `db.t3.xlarge` classes, and gp2/gp3 allocated storage. Unsupported engines, editions, regions, deployment modes, storage types, unknown keys, and mismatched `size`/`instanceClass` fail validation rather than falling back to generic RDS pricing. Non-burstable RDS classes are explicitly unsupported because this bounded regional Express catalog has no verified non-burstable SKU/rate; CWM does not fabricate one or substitute the generic RDS rate. ```json { "id": "rds-sqlserver-1", "type": "database", "name": "SQL Server Express", "provider": "aws", "status": "healthy", "location": { "regionKey": "us-east-1" }, "characteristics": { "serviceFamily": "rds-sqlserver-express", "size": "db.t3.micro", "rdsSqlServerBilling": { "engine": "sqlserver-ex", "region": "us-east-1", "deployment": "single-az", "instanceClass": "db.t3.micro", "storageType": "gp2", "allocatedStorageGiB": 20, "durationHours": 480, "sustainedCpuPercent": 43, "creditMode": "unlimited", "initialCreditBalance": 0, "initialSurplusCreditBalance": 0, "settleSurplusAtEnd": false } } } ``` `sustainedCpuPercent`, duration, storage, region, initial earned-credit balance, and initial surplus-credit balance are caller-provided assumptions. CWM does not infer CPU from low application traffic, connections, AAS, generic simulated database CPU, or SQL statements, and it does not attribute CPU to `RdsAdminService`, SQL Server, the operating system, or another managed-service process. Disabling Unlimited is not modeled and would not remove the underlying CPU demand. One credit is one vCPU-minute. The result separately reports interval `earnedCredits` and `consumedCredits`, the 24-hour earning/debt `creditCap`, `endingCreditBalance`, `endingSurplusCreditBalance`, and `chargedCredits` split into `overflowChargedCredits` plus explicit `settlementChargedCredits`. Low CPU first repays surplus debt before banking earned credits. With `settleSurplusAtEnd: false`, surplus up to the cap remains debt rather than a current charge; set it true only when the interval explicitly assumes stop/delete settlement. The normalized result is returned as `normalizedConfig.resources[].rdsSqlServerBilling` on simulation create/get responses and as `costBreakdown[].rdsSqlServerBilling` on modeled cost output. It separates `computeCost`, `licensingCost`, `storageCost`, `cpuCreditCost`, `totalCost`, and `effectiveHourlyCost`. `licensingCost` is zero because the supported Express rates are license-included rather than because licensing was omitted. `storageType` selects gp2 or gp3 and storage is prorated with a 730-hour reference month. The output also includes its normalized input and an `assumptions` array. Free tier, discounts, tax, backup overage, transfer, additional gp3 IOPS/throughput, and account-specific adjustments are excluded. This is a deterministic constant-CPU interval estimate, not a reconstruction of CloudWatch five-minute samples or AWS hourly settlement. The retained roughly 20-day example is a comparison, not a fitted invoice. With its finalized inputs (480 hours, 43% sustained CPU, zero initial balances, 20 GiB gp2), current published constants produce `$10.56` compute + `$0` separate licensing surcharge + `$1.512328767` storage + `$44.928` charged credits = `$57.000328767` (reported as `$57.000329`). The reported Cost Explorer lines were `$10.30` instance + `$1.49` storage + `$43.16` CPU credits = `$54.95`; the `$2.050329` residual remains disclosed and may reflect the approximate duration/CPU observations, sampling and settlement granularity, or other excluded invoice details. It is not evidence of a CPU root cause. Both authenticated and browser-session right-sizing routes include `rdsSqlServerComparisons` for resources using this contract, even when `hasHint` is false (`GET /simulations/{simulationId}/right-sizing-hint` is the public authenticated path). Each comparison holds absolute vCPU demand and duration constant, starts each alternative with zero earned and surplus balances, returns componentized results, and sets `recommendation` to null. Holding vCPU demand fixed is a comparison assumption, not a prediction that managed-service/background CPU remains identical after resizing. A candidate that cannot represent the fixed demand without exceeding 100% CPU is omitted rather than clamped. The bounded alternatives are only the verified T3 SQL Server Express classes; non-burstable SKUs are unsupported and never fabricated. Treat comparisons as cost evidence, not an architecture choice or migration recommendation. --- ## Pages - [Docs](https://www.cloudworldmodel.ai/docs) - Start with a simulation walkthrough, then explore CWM's MCP agents, RL training environments, chaos experiments, accuracy, fidelity and REST API. - [Compare](https://www.cloudworldmodel.ai/compare) - An honest scope comparison of Cloud World Model with Pinpole, Infracost and Gremlin: what each helps answer, and what simulated results cannot prove. - [September 2026 comparison](https://www.cloudworldmodel.ai/compare/2026-09-29) - A dated, source-linked comparison of CWM, Pinpole, Infracost and Gremlin workflows, with evidence boundaries and a reproducible evaluation checklist. - [Glossary](https://www.cloudworldmodel.ai/glossary) - Definitions of cloud world model, pre-provision simulation, simulated chaos versus production fault injection, and multi-cloud cost and latency simulation. - [Faq](https://www.cloudworldmodel.ai/faq) - Answers about virtual cloud simulations, free credits, model accuracy, chaos experiments and reinforcement-learning environments. - [About](https://www.cloudworldmodel.ai/about) - Why Cloud World Model builds pre-provision multi-cloud simulations for infrastructure teams and AI agents, and how it communicates model limitations. - [Guides/Pre Provision Multi Cloud Simulation](https://www.cloudworldmodel.ai/guides/pre-provision-multi-cloud-simulation) - What pre-provision multi-cloud simulation means, how to check cost and latency evidence, and a valid CWM simulation API example. - [Guides/Rl Autoscaling Environments](https://www.cloudworldmodel.ai/guides/rl-autoscaling-environments) - A grounded guide to training autoscaling policies with CWM's RL environment API, valid action structure and evidence limits. - [Guides/Chaos Without Production](https://www.cloudworldmodel.ai/guides/chaos-without-production) - What simulated chaos can and cannot establish, how it differs from AWS FIS, and a valid CWM chaos API request. - [Home](https://www.cloudworldmodel.ai/) - Platform overview and getting started - [Getting Started](https://www.cloudworldmodel.ai/getting-started) - Quickstart guide for new users; includes interactive 10-step tutorial - [Simulation API Walkthrough](https://www.cloudworldmodel.ai/simulation-walkthrough) - End-to-end developer tutorial chaining all core simulation API calls in sequence: create simulation, step, add a traffic pattern, inject a spike, read metrics history, get right-sizing hints - [Use Cases](https://www.cloudworldmodel.ai/use-cases) - Example use cases for learners and AI agents - [Scenarios](https://www.cloudworldmodel.ai/scenarios) - Pre-built and user-saved simulation scenarios - [Agents](https://www.cloudworldmodel.ai/agents) - RL agent dashboard, API key issuance and revocation, job monitoring - [Pricing](https://www.cloudworldmodel.ai/pricing) - Free tier (1,000 credits/month, no card required), credit packs (Small 10k/$9, Medium 100k/$49, Large 1M/$299), and cost per API call type - [Pricing History](https://www.cloudworldmodel.ai/pricing-history) - Historical and current cloud provider pricing data across all supported providers - [Simulation Fidelity](https://www.cloudworldmodel.ai/fidelity) - Per-provider benchmark sources and stated accuracy ranges: cost ±10%, performance ±15%, calibrated against official AWS, GCP, Azure, OCI, and DigitalOcean pricing pages (June 2026). Lists every covered resource type, instance SKU, hourly rate, and official source URL. - [Supported Regions](https://www.cloudworldmodel.ai/regions) - All cloud regions supported by the simulator across AWS, GCP, Azure, OCI, and DigitalOcean — 177 regions with zone counts, locality type (Availability Zone / Zone / Availability Domain), and provider zone labels - [Benchmark Report](https://www.cloudworldmodel.ai/benchmark) - Full multi-cloud pricing and performance benchmark report comparing AWS, GCP, Azure, OCI, and DigitalOcean across compute, database, storage, and networking tiers - [Simulation Accuracy Benchmark](https://www.cloudworldmodel.ai/accuracy) - Company-owned cwm-bench measurements for the canonical ALB → 2× m5.large EC2 → db.r5.large RDS MySQL Single-AZ architecture in us-east-2 across Idle 10 RPS, Normal 100 RPS, Peak 500 RPS, and Burst 1,000 RPS. The response scores owned app-host CPU and in-VPC internal-LB latency, goodput and CRUD errors against seeded engine predictions; cost is from the AWS us-east-2 price list. Later-day and second-region holdouts have no separate predictions. Data endpoint: GET /api/accuracy-benchmark (public, no auth). - [RL Training Environments](https://www.cloudworldmodel.ai/rl/environments) - Reference guide for training reinforcement learning agents against the cloud simulator: covers the episode lifecycle (reset → step loop → done), the 7-action action space (scale_out, scale_in, adjust_threshold, add_resource, remove_resource, no_op, set_recovery_policy), the 6-field observation vector (rps, cpu_util, instances, traffic, currentTime, tick_seconds), the 5-metric evaluation output (cost_usd_hr, latency_p95, error_rate, uptime, sla_violations), and the reward signal (performance, cost, stability, sla, plus a connection_pressure DB-pool penalty when a database resource is present). Includes runnable multi-episode training-loop examples: Python (`examples/rl_training_loop.py`), JavaScript (`examples/rl_training_loop.js`), Fetch API (`examples/rl_training_loop_fetch.mjs`), TypeScript (`examples/rl_training_loop.ts`), Go (`examples/rl_training_loop.go`, stdlib only — `go run examples/rl_training_loop.go --token $API_KEY`). - [Provider Coverage](https://www.cloudworldmodel.ai/provider-coverage) - Every service, instance family, and pricing model the simulator models for each cloud provider. AWS: 16 EC2 instance types, 16 DB sizes (RDS/Aurora families), ALB, CloudFront, S3, EBS, Lambda, DynamoDB, ElastiCache. GCP: 10 Compute Engine types, Cloud SQL, Spanner, Bigtable, Firestore, Memorystore, Cloud Storage, Persistent Disk, Cloud Run, Cloud Armor. Azure: 10 VM types, Azure SQL (all tiers incl. MI), Cosmos DB, Redis, Blob Storage, Managed Disks, Azure Functions. OCI: 13 compute shapes incl. Bare Metal, MySQL HeatWave, Autonomous DB, Exadata, Block Volume, CDN, WAF. DigitalOcean: 20 Droplet types (Basic, Premium NVMe, CPU-Optimized), Managed DB/Redis/MongoDB/Kafka, DOKS, Spaces. All counts derived from live size arrays; base rates from official pricing constants (June 2026). Coverage snapshot version updated in COVERAGE_VERSION and validated against a 90-day freshness window in CI. - [Privacy Policy](https://www.cloudworldmodel.ai/privacy) - Privacy policy - [Terms of Service](https://www.cloudworldmodel.ai/terms) - Terms of service - [Blog](https://www.cloudworldmodel.ai/blog) - Insights from building Cloud World Model — product launches, engineering decisions, and what we learned along the way. - [Blog/What We Learned Product Hunt](https://www.cloudworldmodel.ai/blog/what-we-learned-product-hunt) - Retrospective on the Product Hunt launch: key takeaways from finishing #5 Product of the Day and 100+ user comments - [Blog/X402 Cloud Simulation Api](https://www.cloudworldmodel.ai/blog/x402-cloud-simulation-api) - How we built native x402 payable API support so AI agents can run simulations and pay per request — no human-in-the-loop billing setup required. - [Blog/Owned Aws Accuracy Benchmark](https://www.cloudworldmodel.ai/blog/owned-aws-accuracy-benchmark) - Earlier campaign methodology; the live accuracy page now scores owned app CPU and in-VPC latency rather than documentation estimates. - [Blog/Cloud Simulation Step Internals](https://www.cloudworldmodel.ai/blog/cloud-simulation-step-internals) - A technical walkthrough of the step-hybrid engine: request body, state transitions, metric outputs, and ML/rules blending to a final recommendation. - [Blog/Cloud Pricing Trends August 2026](https://www.cloudworldmodel.ai/blog/cloud-pricing-trends-august-2026) - We verify cloud pricing monthly across AWS, GCP, Azure, OCI, and DigitalOcean. Here's what changed in August 2026 and what it means for your budget. - [Blog/Mcp Server Launch](https://www.cloudworldmodel.ai/blog/mcp-server-launch) - Connect any MCP-compatible AI client to the Cloud World Model simulation engine with a single URL — no install. Anonymous demo mode, no key required. - [MCP Server Quick Start](https://www.cloudworldmodel.ai/mcp-quickstart) - Connect to the Cloud World Model MCP server in minutes via Smithery or the hosted HTTP endpoint — no local install required --- ## API Surface All external APIs are served under the `/api` prefix. Simulation creation and key management endpoints are public (no auth required for demo purposes). RL, chaos, multi-cloud, prediction, optimization, and AI analysis endpoints require a Bearer API key or an x402 payment header. Base URL: `https://www.cloudworldmodel.ai/api` ### x402 Pay-Per-Request — Discovery Summary Cloud World Model is an **x402-compatible seller** (x402 is a Linux Foundation open protocol). AI agents can call any metered endpoint without an account or API key by selecting an advertised USDC payment option and attaching its signed payload header. Base is always the primary option; Solana is an optional second option when its wallet is configured. | Field | Value | |---|---| | Discovery endpoint | `GET /api/billing/x402/config` | | Base option | `eip155:8453` (Base, Ethereum L2), USDC, header `X-PAYMENT` | | Optional Solana option | `solana:5eykt4UsFv8P8NJdTREpY1vzqKqZKvdp`, USDC-SPL, header `PAYMENT-SIGNATURE` | | No signup required | Yes — no account, no API key needed for x402 callers | | Balance lookup | `GET /api/billing/x402/balance?address=0x...` | | Transaction history | `GET /api/billing/x402/transactions?address=0x...&limit=50` | | x402scan listing | https://www.x402scan.com/server/1a8052f0-c764-420a-96a5-f50d3c696795 | Call `GET /api/billing/x402/config` first. Its `paymentOptions` array is authoritative: select an option, use its exact `network`, `asset`, `payTo`, and `header`, and include the Solana option's `feePayer` when present. The legacy Base-primary fields and optional `solana` object remain available for existing clients. The endpoint also provides exact USDC amounts per call type, the facilitator URL, and the credits-per-USDC conversion rate; it is public and requires no authentication. **Per-use-case cost estimates** (based on current $0.0010/step and $0.0010/AI-analysis pricing; analysis jobs cost $0.0050): | Use case | Calls | Estimated cost | |---|---|---| | 100-step RL training loop | 100 × `rl.step` | ≈ $0.10 | | 1 000-step RL training run | 1 000 × `rl.step` | ≈ $1.00 | | AI explanation (explain) | 1 × `ai.explain` | $0.0010 | | AI optimization suggestions | 1 × `ai.optimize` | $0.0010 | | AI troubleshooting guide | 1 × `ai.troubleshoot` | $0.0010 | | AI bottleneck analysis | 1 × `ai_bottleneck` | $0.0010 | | Single chaos scenario | 1 × `chaos.run` | $0.0050 | | Batch of 20 chaos scenarios | 20 × `chaos.run` | $0.10 | | Multi-cloud cost comparison | 1 × `multicloud.explore` | $0.0050 | | Predictive scaling validation | 1 × `prediction.validate` | $0.0050 | | Hybrid simulation step | 1 × `simulation_step_hybrid` | $0.0010 | Always fetch `GET /api/billing/x402/config` for the authoritative live price table — the values above reflect current pricing but the config endpoint is the single source of truth. **Minimal Python agent snippet** (`httpx-x402` library): ```python import httpx from x402.httpx import wrap_httpx_client from x402.types import X402Config # 1. Discover live prices (no auth required) config = httpx.get("https://www.cloudworldmodel.ai/api/billing/x402/config").json() # 2. Wrap your httpx client with x402 auto-pay support client = wrap_httpx_client( httpx.Client(), private_key="0xYOUR_PRIVATE_KEY", # EVM wallet key — keep secret ) # 3. Call a metered endpoint — payment happens automatically on 402 response = client.post( "https://www.cloudworldmodel.ai/api/rl/environments/{env_id}/step", headers={"Content-Type": "application/json"}, json={"action": "scale_up"}, ) print(response.json()) # observation, reward, done ``` ### Authentication & Key Management | Method | Path | Auth | Description | |---|---|---|---| | POST | /keys | Public | Create a new API key; the raw key is returned once - store it securely | | DELETE | /keys/{keyId} | Public | Revoke an API key | **Key scopes**: `read`, `write`, `admin`. Default: `[read, write]`. ### Wallet Auth (EIP-191 sign-in — no API key required) Agents with an EVM wallet can authenticate without creating a traditional API key. The two-step flow issues a 24-hour JWT session token that can be used as a Bearer token on `GET /simulations` and `GET /simulations/{simulationId}`. **Flow:** call `POST /wallet-auth/challenge` with your wallet address → sign the returned `message` with `personal_sign` (EIP-191) → call `POST /wallet-auth/verify` with the signature → receive a `token` → use `Authorization: Bearer ` on simulation endpoints. | Method | Path | Auth | Description | |---|---|---|---| | POST | /wallet-auth/challenge | Public | Request an EIP-191 sign-in challenge (nonce + message to sign). Rate-limited: 10/min per IP. | | POST | /wallet-auth/verify | Public | Submit the EIP-191 signature to receive a 24-hour JWT session token. | | POST | /wallet-auth/revoke | Public | Immediately revoke a wallet session JWT by recording its jti in the server-side revocation set. Pass the token in the request body; no Authorization header needed. Revocation is in-memory and does not survive a server restart. Rate-limited: 10/min per IP. | | POST | /wallet-auth/x402-session | x402 micro-payment | Issue a 24-hour wallet session JWT via USDC micro-payment on Base (x402 protocol). Idempotent per transaction hash. Token validity survives restarts only when WALLET_AUTH_JWT_SECRET is set. | | POST | /wallet-auth/revoke | Public | Immediately revoke a wallet session JWT before its natural 24-hour expiry. Pass `{ "token": "" }` in the body. Revoked tokens are rejected with 401 on all subsequent requests. Revocation is in-memory and does not survive a server restart. | **Wallet session scope:** Read-only (`read`). The session can list and retrieve simulations that were claimed by the wallet address (`ownerWallet` field). `GET /simulations` returns a paginated object `{ simulations, total, limit, offset }` for wallet sessions (rather than a plain array for API-key sessions). Mutations (`POST`, `PATCH`, `DELETE`) still require an API key. ### Simulation API Create and manage simulations, step them forward, inject traffic or failures, read metrics and events, and call AI-powered analysis helpers. **Lifecycle** | Method | Path | Auth | Description | |---|---|---|---| | GET | /prediction/generic-shapes | Bearer (read) | Read-only, resolver-derived inventory of named compute catalog gaps and conditional generic fallback categories. Unknown/custom compute is a wildcard. Per-simulation reasons and `legacyGeneric` flags appear in versioned `predictionEvidence`; the inventory is not a performance or pricing reference. | | GET | /ui/prediction/generic-shapes | Browser UI (rate-limited) | Same generic-shape inventory for Workspace and Accuracy methodology readers; no simulation is created or mutated. | | POST | /simulations | Bearer (write) or x402 (`simulation_create`, $0.0010) | Create a simulation with either a live `scenarioId` or an explicit cloud resource graph (compute, database, storage, networking). | | GET | /simulations | Bearer (read) or x402 (`simulation_list`, $0.0010) | List all simulations owned by the authenticated key | | GET | /simulations/{simulationId} | Bearer (read) or x402 (`simulation_get`, $0.0010) | Get full simulation state (resources, time step, traffic, autoscaling history) | | GET | /simulations/{simulationId}/cost-breakdown | Bearer (read) or x402 (`simulation_cost_breakdown`, $0.0010) | Latest per-resource hourly cost with raw status plus availabilityState/isRoutable/routedRps — distinguish degraded-but-serving from unavailable residual billing | | POST | /simulations/{simulationId}/provider-api-limits | Bearer (read) or x402 (`provider_api_limits`, $0.0010) | Run a bounded synchronous provider API-limit simulation for an owned simulation. One service per request: top-level `service` applies to all operation groups. AWS catalog selectors: `ec2` / `mutating` / `compute` (`aws.ec2.mutating-api`); `elbv1` or `elbv2` / `resource-intensive`, `registration`, `non-mutating`, or `mutating` / `network` (`aws.elbv1.resource-intensive`, `aws.elbv1.registration`, `aws.elbv1.non-mutating`, `aws.elbv1.mutating`, `aws.elbv2.resource-intensive`, `aws.elbv2.registration`, `aws.elbv2.non-mutating`, `aws.elbv2.mutating`). ELB v1/v2 buckets are distinct; do not infer an operation category for uncategorized actions. Aurora Serverless v1 Data API: `rds-data-api-aurora-serverless-v1` / `database` / `requests-per-second` (`aws.rds-data-api-aurora-serverless-v1.requests-per-second`, 1,000 req/s per account and Region) OR `concurrent-requests` (`aws.rds-data-api-aurora-serverless-v1.concurrent-requests`, 500 concurrent for ONE cluster using the SAME secret; overflow queues). These are independent models, NOT joint enforcement of the same workload. Both require `quotaScope.account`; concurrency also requires `quotaScope.resource` as an opaque cluster-secret-pair label, never the actual secret. No documented burst for the rate model; the fixed 1-second scheduler window approximates the AWS limit, not AWS timing. Source: https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/CHAP_Limits.html. Azure Resource Manager read/write also has defaults. IAM, generic RDS control-plane, Aurora Serverless v2 Data API, provisioned Aurora Data API, EC2 Auto Scaling, Application Auto Scaling, and S3 control-plane lack a numeric public default here; S3 object per-prefix throughput is data-plane guidance. Supported tuples resolve as `catalog_default` and return provenance; undocumented tuples return `unsupported_policy` without a scoped caller override. Models throttling, queueing, throughput, completion, retries, timeouts, failures, and concurrency. Above 1,000 planned units the response uses bounded `summary_only` mode with no per-operation records; sweep comparisons retain reconciled planned/attempted/completed/failed/throttled/retry/pending/terminal counters and queue/timing fields, and `recommendedConcurrency` is null unless a candidate fully completes within the throttle constraint. A concurrency sweep shares a 1,000,000-event and 100,000-unit aggregate budget across candidates, the whole request has a 2-second wall-clock deadline, and two executions may run concurrently. A third request returns non-retryable 409 `PROVIDER_LIMIT_BUSY`; an identical active or interrupted request returns non-retryable 409 `PROVIDER_LIMIT_DUPLICATE` with a five-minute duplicate guard. Does not call a provider, discover live quotas, change /step state, or act as a provider SDK. GCP, DigitalOcean, and OCI tuples without a scoped override return `unsupported_policy`. | | PATCH | /simulations/{simulationId} | Bearer (write) | Partially update simulation properties (name, description, traffic, resources) | | DELETE | /simulations/{simulationId} | Bearer (write) | Delete a simulation and all associated data | | POST | /simulations/{simulationId}/claim | Bearer (write) | Claim ownership of an unowned demo simulation; returns 200 on success (idempotent if caller already owns it), 409 if owned by another key | **ARO cost-breakdown unit note:** compare like-for-like units. The ARO `publishedUnitRatePerHour` is `$0.171/hour per 4 vCPU`; for three 2-vCPU workers, `effectiveWorkerVcpu` is 6 and the topology-derived `derivedPlatformFeePerHour` / platform-row `costPerHour` is `$0.2565/hour` (`$0.171 × 6 ÷ 4`). This is not pricing drift. Platform, estimated control-plane, and worker charges are separate additive rows. **Aurora Serverless v1 Data API provider-limit requests:** A request mixing the `requests-per-second` and `concurrent-requests` selectors is rejected. Submit separate requests; these are independent models, not simultaneous enforcement of the same workload. **Ownership & expiry — read this before reusing simulation IDs** Simulations come in two flavours and the difference matters for any agent or tool that stores IDs: - **Owned (API-created) simulations are persistent.** A simulation created through `POST /api/simulations` with a Bearer token — or one you `claim` — is bound to your API key and **never expires**. These are the only simulations reachable through the authenticated `/api/simulations/*` API (including the `validate-*` endpoints). Treat the returned ID as the durable handle you manage. - **Unowned (browser-workspace / demo) simulations are ephemeral.** Simulations built interactively in the browser workspace have no API-key owner. They live only behind the session-cookie UI routes, are **never** reachable through the authenticated API, and **auto-expire after ~5 minutes**. Their IDs are not durable. **Claim-before-validate workflow.** To persist and then validate a simulation you built in the browser workspace, claim it *before it expires*: `POST /api/simulations/{simulationId}/claim` with a write-scoped Bearer token binds it to your key (idempotent if you already own it; `409` if another key owns it). Once owned, it stops expiring and the `validate-*` endpoints work on it. **Guidance for agents/tools.** Always **create (or claim) simulations through the API** and treat the returned ID as the thing you manage. **Do not reuse IDs from earlier browser/demo runs** — an unowned ID from a previous session has almost certainly expired. Hitting `GET /api/simulations/{id}` or any `validate-*` endpoint with such an ID returns `404 {"error":"Simulation not found or expired", "reason": "...", "remedy": "..."}`. The `remedy` field tells you to create an owned simulation via `POST /api/simulations` (or claim an unowned one). A `404` never confirms whether a simulation owned by a *different* key exists — a sim owned by another key returns `403 {"error":"Access denied"}` instead. **`validate-*` works on any owned simulation, of any complexity.** The cost/performance/combined validators map each resource to the nearest pricing & performance benchmark by provider + service family + size, so they work on hybrid, multi-provider, autoscaled, and post-chaos simulations — not just simple single-resource "fidelity match" sims. You do not need a special benchmark-shaped simulation to call them; any owned simulation you can `GET` can be validated. **Kubernetes cluster size — node-pool characteristics.** A resource with `type: "kubernetes"` models a managed container node pool (EKS / GKE / AKS / OKE / DOKS). Its size is set through three `characteristics` fields: | Field | Type | Description | |---|---|---| | `nodeCount` | number? | Current number of worker nodes in the node pool. The cluster's `maxThroughput` (serving capacity) and hourly cost scale proportionally with this count. | | `minNodes` | number? | Minimum worker nodes the pool will scale in to. Defaults to the provider's autoscaling profile minimum when omitted. | | `maxNodes` | number? | Maximum worker nodes the pool will scale out to. Defaults to the provider's autoscaling profile maximum when omitted. | During `/step` the autoscaler adds or removes worker nodes between `minNodes` and `maxNodes`, updating `nodeCount` in place and recomputing the cluster's `maxThroughput` and hourly cost (`perNodeRate × nodeCount`). Set `nodeCount` to size the cluster at creation; set `minNodes`/`maxNodes` to bound how far it can autoscale. **Multiple node pools per cluster.** For clusters with heterogeneous workloads, set `characteristics.nodePools` to an array of pool objects instead of the single-pool fields above. Each pool scales independently within its own bounds and is billed at its own per-node rate: | Field | Type | Description | |---|---|---| | `name` | string? | Human-readable node pool name (e.g. `general`, `gpu`). | | `nodeCount` | number? | Current number of worker nodes in this pool. Scales this pool's serving capacity and cost proportionally. | | `minNodes` | number? | Minimum worker nodes this pool will scale in to. Defaults to the provider's autoscaling profile minimum when omitted. | | `maxNodes` | number? | Maximum worker nodes this pool will scale out to. Defaults to the provider's autoscaling profile maximum when omitted. | | `perNodeRate` | number? | Hourly cost per node for this pool. Falls back to the provider default rate when omitted. | When `nodePools` is present, the legacy single-pool fields (`nodeCount`, `minNodes`, `maxNodes`) are ignored; when it is absent, the cluster behaves as a single pool described above. **Resource capacity tuning — `characteristics` fields.** Beyond the Kubernetes node-pool fields above, the `characteristics` object exposes several capacity knobs the simulation engine uses to model how a resource behaves under load. All are optional; sensible defaults apply when omitted. | Field | Type | Applies to | Description | |---|---|---|---| | `maxConnections` | number? | database | Connection-pool capacity (max concurrent connections). The engine models pool pressure as `activeConnections / maxConnections` (surfaced as `metrics.connection_pressure`); raising it gives the DB more headroom before saturation. Defaults to 100. | | `capacityGB` | number? | storage | Provisioned storage capacity in gigabytes. Used with `maxIops` to model block-storage throughput and IOPS utilization. | | `maxIops` | number? | storage | Provisioned maximum IOPS for a block-storage volume. The engine models IOPS utilization against this ceiling (surfaced as `metrics.storageIopsUtilization`); falls back to the provider's default volume profile IOPS when omitted. | | `cacheHitRate` | number? | cache / CDN | Expected cache hit rate as a fraction between 0 and 1 (e.g. `0.85` = 85%). Higher values reduce load reaching downstream origin/database resources and lower effective latency. Defaults to 0.8 for cache resources. | | `openshiftOffering` | `rosa-hcp` \| `rosa-classic` \| `aro` \| `openshift-dedicated` \| `self-managed`? | kubernetes | Optional OpenShift distribution overlay. The backing `provider` still determines regions, network behavior, worker infrastructure, and provider latency. ROSA HCP/Classic require AWS, ARO requires Azure, OpenShift Dedicated supports AWS/GCP, and self-managed supports existing CWM providers. Cost breakdowns keep worker, platform, and control-plane charges separate. ROSA's published worker service fee is $0.171 per 4 vCPU-hour across supported AWS standard regions; ROSA HCP adds a published $0.25 per cluster-hour fee. ARO's canonical East US D4s v3 (4 vCPU) OpenShift license line is $124.830/month ($0.171/hour using 730 hours/month), verified 2026-08-27 from https://azure.microsoft.com/en-us/pricing/details/openshift/. It is marked official only when `location.regionKey` identifies Azure East US (including the stored `eus` shortcode); other or missing regions use the same amount as an estimated reference assumption. OpenShift Dedicated and self-managed subscription allocations remain estimated where no comparable universal hourly line exists. Omit for generic Kubernetes. | | `autoscaling` | boolean? | compute | Marks a compute resource as the autoscaling primary. When multiple compute resources exist, the engine targets the one flagged `autoscaling: true`; otherwise it falls back to the first compute resource. | **Simulation Steps** | Method | Path | Auth | Description | |---|---|---|---| | POST | /simulations/stateless | Bearer (write) or x402 (`simulation.stateless`, $0.0010) | Run up to 10 deterministic in-memory production-engine steps and receive synchronous cost, performance, utilization, error, reliability, coverage, and an autoscaling audit. `autoscaling.minInstances` is materialized before the first step; the audit reports requested/effective bounds, CPU target, enabled state, and initial fleet size. Creates no simulation or job records. REST-only: MCP cannot preserve the HTTP x402 challenge, settlement, and replay boundary. | | POST | /simulations/{simulationId}/step | Optional | Advance simulation by one time step; applies patterns, capacity model, and autoscaler | | POST | /simulations/{simulationId}/step-hybrid | Bearer (write) or x402 (`simulation_step_hybrid`, $0.0010) | Advance using the Hybrid (rule + ML) Prediction Engine; returns blending metadata | | GET | /simulations/{simulationId}/hybrid-result | Bearer (read) | Retrieve cumulative hybrid decision history and summary statistics | **`/step` response — `coverageSummary` field** Every `/step` response includes a `coverageSummary` object that tells you what the simulator could and could not model for this resource mix: | Field | Type | Description | |---|---|---| | `level` | `"full" \| "partial" \| "limited"` | Resource/behavior and base-rate coverage level only. `"full"` does not mean complete economic or invoice coverage. | | `modeledCount` | integer | Resources whose cost and behaviour are fully modeled by deterministic provider-calibrated rules. | | `estimatedCount` | integer | Resources where values are estimated rather than exact. | | `knownGapCount` | integer | Resources or cost categories that are known gaps (present in the topology but not modeled). | | `notObservableCategories` | string[] | Cost categories that cannot be observed regardless of resource configuration (e.g. `["Support plan costs"]`). When non-empty, `level` is always `"limited"`. | | `billingPolicyCoverage` | object | Separate economic-policy signal: `{ complete, status, exclusions }`. Each RDS resource exposes an unpriced `RDS Extended Support` exclusion (`priced: false`); no surcharge is estimated. | | `resources` | object[] | Per-resource coverage detail. Each entry has: `name`, `resourceType` (`"cost_category"` for category rows), `costFidelity` (`modeled \| estimated \| known_gap \| not_observable`), `behaviourFidelity` (same enum), `shareOfSimulatedCost` (0–1, 0 for gaps), `reason` (human-readable explanation), `coverageBasis` (`deterministic \| ml \| estimated \| unsupported`). | Use `level` only for resource/behavior fidelity, `resources[].costFidelity` for per-resource drill-down, and `billingPolicyCoverage.complete` for economic-policy completeness. Never interpret `level: "full"` as “the whole bill is covered.” Important: `costFidelity: "modeled"` means a resource's cost is backed by deterministic CWM rules using published list prices — modeled does not mean billing-validated. Real invoices can differ due to committed-use discounts, support surcharges, marketplace licensing, and on-demand price changes regardless of fidelity level or `modeledCostShareOfTotal`. **step-hybrid request body** ```json { "trafficRPS": 800, "config": { "blendingWeights": { "rules": 0.5, "ml": 0.5 }, "confidenceThreshold": 0.7, "safetyBounds": { "maxLatencyDeviation": 50, "maxCostDeviation": 0.2, "maxErrorRateDeviation": 5 }, "fallbackToRules": true } } ``` `blendingWeights` controls the initial rule/ML split. The effective ML weight is scaled by the model's reported confidence: `effectiveMlWeight = normalizedMlWeight × confidence`. When `effectiveMlWeight` would fall below `confidenceThreshold` (or when any blended field deviates from the rule baseline by more than a safety bound), `fallbackToRules: true` lets the referee override blending entirely. The result is recorded in `hybridDecision.blendingApplied` (boolean) and explained in `hybridDecision.explanation` (string). **step-hybrid DigitalOcean quickstart (JavaScript)** ```javascript // 1. Create a DigitalOcean simulation with an s-2vcpu-4gb Droplet const createRes = await fetch("https://www.cloudworldmodel.ai/api/simulations", { method: "POST", headers: { "Authorization": "Bearer your_api_key_here", "Content-Type": "application/json" }, body: JSON.stringify({ name: "do-droplet-hybrid", resources: [{ id: "web-1", type: "compute", name: "Web Server", provider: "digitalocean", config: { instanceType: "s-2vcpu-4gb" }, }], }), }); const sim = await createRes.json(); // 2. Run step-hybrid — referee blends the deterministic rule step with the ML model const hybridRes = await fetch( `https://www.cloudworldmodel.ai/api/simulations/${sim.id}/step-hybrid`, { method: "POST", headers: { "Authorization": "Bearer your_api_key_here", "Content-Type": "application/json" }, body: JSON.stringify({ trafficRPS: 800, config: { blendingWeights: { rules: 0.5, ml: 0.5 }, confidenceThreshold: 0.7, fallbackToRules: true, }, }), } ); const result = await hybridRes.json(); // 3. Read the blended metrics and the referee's decision const { metrics, hybridDecision } = result; console.log("P95 latency :", metrics.latencyP95, "ms"); console.log("Cost/hr : $" + metrics.costPerHour.toFixed(4)); console.log("Blending :", hybridDecision.blendingApplied); // true → ML output was trusted; effective ML weight = 0.5 × ML confidence // false → referee fell back to rules only (low confidence or large deviation) console.log("Explanation :", hybridDecision.explanation); ``` **step-hybrid DigitalOcean quickstart (Python)** ```python import requests API_KEY = "your_api_key_here" BASE = "https://www.cloudworldmodel.ai/api" headers = {"Authorization": f"Bearer {API_KEY}", "Content-Type": "application/json"} # 1. Create a DigitalOcean simulation with an s-2vcpu-4gb Droplet sim = requests.post(f"{BASE}/simulations", headers=headers, json={ "name": "do-droplet-hybrid", "resources": [{ "id": "web-1", "type": "compute", "name": "Web Server", "provider": "digitalocean", "config": {"instanceType": "s-2vcpu-4gb"}, }], }).json() # 2. Run step-hybrid — referee blends the deterministic rule step with the ML model result = requests.post( f"{BASE}/simulations/{sim['id']}/step-hybrid", headers=headers, json={ "trafficRPS": 800, "config": { "blendingWeights": {"rules": 0.5, "ml": 0.5}, "confidenceThreshold": 0.7, "fallbackToRules": True, }, }, ).json() # 3. Read the blended metrics and the referee's decision metrics = result["metrics"] hybrid = result["hybridDecision"] print(f"P95 latency : {metrics['latencyP95']:.1f} ms") print(f"Cost/hr : ${metrics['costPerHour']:.4f}") print(f"Blending : {hybrid['blendingApplied']}") # True → ML output was trusted; effective ML weight = 0.5 × ML confidence # False → referee fell back to rules only (low confidence or large deviation) print(f"Explanation : {hybrid['explanation']}") ``` Key `hybridDecision` fields returned by `step-hybrid`: | Field | Type | Description | |---|---|---| | `blendingApplied` | boolean | `true` when the ML output was blended into the final metrics; `false` when the referee used rules only | | `explanation` | string | Human-readable summary of which path the referee took and why (confidence level, deviation check, override reason) | | `hybridMetrics` | object | Pre-failure blended snapshot — same shape as the top-level `metrics` but captured before active failure effects are applied | | `mlConfidence` | number | Raw confidence score reported by the ML model (0–1); scaled by the normalized ML weight to produce the effective ML contribution | The blending API is identical across all five providers — swap `provider` and `instanceType` to use AWS, GCP, Azure, or OCI. The `hybridDecision` object and its fields are always present in the response regardless of which provider the simulation uses. **Step response — metrics fields** Every `/step` response includes a `metrics` object. Optional fields appear only when the relevant resource types are present: | Field | Type | Description | |---|---|---| | `latencyP50` | number | Median response latency (ms) | | `latencyP95` | number | 95th-percentile latency (ms) | | `latencyP99` | number | 99th-percentile latency (ms) | | `cpuUsage` | number | Average CPU utilization across compute resources (%) | | `memoryUsage` | number | Estimated memory utilization (%) | | `throughput` | number | Effective requests per second (goodput after errors) | | `errorRate` | number | Error rate (%) | | `costPerHour` | number | Estimated infrastructure cost (USD/hr) | | `cacheHitRate` | number? | Cache hit rate (%) — only when cache resources exist | | `queueDepth` | number? | Pending messages in queue — only when queue resources exist | | `k8sNodeUtilization` | number? | Kubernetes node CPU utilization (%) — only when k8s resources exist | | `storageIopsUtilization` | number? | OCI Block Volume IOPS utilization (%) — only when OCI block storage exists | | `connectionPressure` | number? | DB connection-pool pressure: `activeConnections / maxConnections`, capped at 3.0. Only present when database resources exist. Values > 1.0 indicate pool exhaustion. Use this to tune agent reward functions based on connection saturation margin. | **Error breakdown and serving CPU overload** `metrics.errorBreakdown` (and the snake_case `error_breakdown` returned by RL endpoints) reports pre-clamp error contributors in percentage-point units. `queueAbsorption` is a positive reduction; the other contributors are additive before the overall error-rate cap. `error_rate` in RL responses is a fraction, but the breakdown values remain percentage points, matching `error_rate × 100`. `cpuOverload` is specifically serving-CPU telemetry, not a compute-failure count. It uses only active request-serving compute and Kubernetes resources that are actually serving traffic. Parked or unavailable nodes are excluded from both the numerator and denominator. Database and network resource CPU is separate: 100% CPU on a database or load balancer does not increase `cpuOverload`; inspect `poolSaturation`/`dbFailure` and the relevant network or latency signals instead. For each eligible serving resource, the CPU excess is `max(0, (cpuUsagePercent - 80) / 20)`. The API averages those per-resource excesses and multiplies the result by the provider slope. Ordinary/custom AWS uses a 10 percentage-point slope: one eligible AWS node at 100% CPU contributes 10 points. For example, if two serving AWS nodes are at 100% and 70%, `cpuOverload` is 5 points. A third node at 100% that is parked/unavailable, a database at 100%, and a network resource at 100% do not change that 5-point value because none is an eligible serving compute node. The owned AWS CRUD benchmark can intentionally suppress this contributor and report zero as an internal calibration policy because that campaign measured no separate CPU-error bucket. That is benchmark behavior, not a generic API rule; ordinary AWS simulations still use the 10-point serving-CPU slope. Example (partial fleet failure with dependency CPU kept separate): ```json { "metrics": { "errorBreakdown": { "computeFailure": 0, "cpuOverload": 5, "dbFailure": 0, "poolSaturation": 0 } }, "interpretation": "web-1=100% and web-2=70% are serving; web-3=100% is parked/unavailable; db-1=100% and alb-1=100% are database/network CPU and are excluded from cpuOverload" } ``` **Step response — database_overload event metadata** When a `database_overload` failure is active, the generated event includes a `metadata` field with structured connection arithmetic: ```json { "severity": "error", "message": "Database overload: my-db is critically overloaded - modeled query-service contention at 300 RPS", "resource": "my-db", "metadata": { "activeConnections": 60, "maxConnections": 800, "connectionUtilizationPct": 7.5, "offeredDbRps": 300, "modeledServiceCapacityRps": 120 } } ``` `activeConnections` is modeled from offered DB traffic (0.2 connections per RPS, divided across databases), not synthesized from severity. `connectionUtilizationPct` compares that demand to the effective connection budget; overload can reflect query-service contention even when the pool is *not* saturated. `modeledServiceCapacityRps` is a simulator assumption, not a measured AWS limit. RL agents should distinguish this incident signal from actual `connectionPressure` in step metrics. --- ## Resilience Contracts — Retry Amplification & Circuit-Breaker Modeling CWM can model retry amplification, circuit-breaker state, rate-limiting, and cascading-failure depth across declared dependency edges. Enable it by attaching a `resilienceConfig` at simulation creation (`POST /simulations`) or update it any time via `PATCH /simulations/{simulationId}`. ### `resilienceConfig` schema (attached to a Simulation) ```json { "resilienceConfig": { "enabled": true, "version": 1, "dependencies": [ { "id": "web-to-db", "sourceId": "web-1", "targetId": "db-1", "requestRatio": 1, "retryPolicy": { "maxRetries": 3, "backoffMs": 100, "backoffMultiplier": 2, "jitterRatio": 0.1, "timeoutMs": 2000, "retryBudgetRatio": 2, "retryBudgetRps": 1000 }, "protection": { "circuitBreaker": { "enabled": true, "failureRateThreshold": 0.5, "minimumRequests": 20, "openSteps": 3, "halfOpenMaxRequests": 10 }, "rateLimitRps": 5000, "loadShedding": false } } ], "scheduledFaults": [ { "id": "db-capacity-fault", "type": "capacity_limit", "targetResourceId": "db-1", "startStep": 5, "endStep": 15, "capacityPercent": 30 } ], "maxCascadeDepth": 4, "maxGeneratedRps": 100000, "retryGeneratedTrafficAffectsCost": false } } ``` **Bounded safety limits** (server-enforced — any config exceeding these is rejected with 400): - `dependencies`: max 64 edges - `scheduledFaults`: max 32 faults - `maxRetries` per dependency: max 8 - `maxCascadeDepth`: max 8 - `maxGeneratedRps`: max 500,000 combined retry-generated and auth/token RPS - `maxStepWork`: max 2,048 dependency-path work items per step **`scheduledFault.type` values:** `capacity_limit`, `concurrency_limit`, `latency`, `error_rate`. Each type uses a different set of numeric fields; the server validates consistency. ### Resilience telemetry in step responses When `resilienceConfig.enabled` is true, every `/step` and `/metrics` response gains two additional fields: | Field | Type | Description | |---|---|---| | `retryAmplificationFactor` | number \| null | Headline metric: `totalAttemptedRps / originalClientRps`. Values > 1.0 mean retries are generating more traffic than original requests. `null` = model ran but traffic was zero. **Absent** (not in response) = resilience model disabled. Do NOT synthesize 1.0 for absent — that would incorrectly imply modeling occurred. | | `resilience` | object | Full `ResilienceTelemetry` block — simulation-level totals + per-dependency `paths[]` array. Each path carries `circuitState` (`closed|open|half_open`), `availability`, `retryRps`, `shedRps`, `authTokenRps`, `authTokenFailedRps`, and `saturated`. Use `retryAmplificationFactor` as the headline; drill into `resilience.paths[]` for per-edge detail. | The compact MCP step response (`simulation.step`, `simulation.metrics`) also surfaces: - `retryAmplificationFactor` (same semantics as above) - `resilienceDiagnostics` — bounded summary `{incidentOutcome, pathCount, bounded}` (paths array omitted for token efficiency); `bounded: true` means a generated-traffic, traversal-work, or cascade-depth limit truncated model work ### `POST /simulations/{simulationId}/resilience/compare` Side-by-side replay of two resilience configurations across the simulation's current traffic profile. **Auth:** Bearer token with `write` scope, or wallet-session JWT from an existing x402-paid workflow. No additional x402 payment is charged. **Request body:** | Field | Type | Default | Description | |---|---|---|---| | `steps` | integer | 20 | Steps to replay (1–120). More steps = more representative peak metrics. | | `baselineConfig` | ResilienceConfig? | simulation's current resilienceConfig | Explicit baseline config. Returns 400 when both this and the simulation's resilienceConfig are absent. | | `mitigatedConfig` | ResilienceConfig | required | The mitigated config to evaluate (e.g. with circuitBreaker.enabled: true). | | `mitigatedResources` | Resource[]? | simulation's resources | Alternative resource array for the mitigated run. | | `mitigatedAutoscalingConfig` | object? | — | Autoscaling config override for the mitigated run. | **Response: `ResilienceComparison`** ```json { "seed": 42, "startStep": 10, "traffic": 5000, "steps": 20, "baseline": { "peakRetryAmplificationFactor": 1.87, "peakErrorRate": 4.2, "peakLatencyP95": 340, "totalServedRequests": 195000, "totalShedRequests": 5000, "finalOutcome": "degraded" }, "mitigated": { "peakRetryAmplificationFactor": 1.42, "peakErrorRate": 2.4, "peakLatencyP95": 280, "totalServedRequests": 198000, "totalShedRequests": 2000, "finalOutcome": "protected" }, "delta": { "retryAmplificationFactor": -0.45, "errorRate": -1.8, "latencyP95": -60, "shedRequests": -3000 } } ``` **Reading the delta:** Negative values indicate improvement in the mitigated run. `delta.retryAmplificationFactor < 0` means fewer retry storms; `delta.errorRate < 0` means fewer errors; `delta.shedRequests < 0` means less load-shedding. **`incidentOutcome` values:** `stable` | `degraded` | `cascading` | `protected` | `recovered` ### Resilience comparison MCP workflow (agent checklist) 1. **Configure resilience** — create simulation via `simulation.create` with `resilienceConfig` (or use `simulation.update` to attach/replace it). 2. **Step several times** — call `simulation.step` 10–20 times to establish a traffic baseline. Check `retryAmplificationFactor` in each response. 3. **Compare configs** — call `simulation.compare_resilience` with the current config as baseline and a modified config (e.g. `circuitBreaker.enabled: true`, or lower `retryBudgetRatio`) as mitigated. 4. **Interpret the delta** — if `delta.retryAmplificationFactor` is negative and `mitigated.finalOutcome` is `protected` (not `cascading`), the mitigation is effective. 5. **Iterate** — if the outcome is still `degraded` or `cascading`, try: (a) tightening `circuitBreaker.failureRateThreshold`, (b) lowering `retryBudgetRatio`, (c) enabling `loadShedding`, (d) reducing `maxCascadeDepth`. 6. **Persist** — once the delta is satisfactory, call `simulation.update` with the mitigated config to permanently replace the baseline. 7. **Next tool** — continue stepping with `simulation.step` or call `simulation.metrics` to see resilience telemetry over the full run history. --- **Traffic & Failure Injection** | Method | Path | Auth | Description | |---|---|---|---| | POST | /simulations/{simulationId}/inject-traffic | x402 \| Bearer (write) | Inject traffic — provide `targetRps` (absolute RPS) or `deltaPercent` (relative change, e.g. 75 for +75%); pass `{"random": true}` to opt into an uncontrolled random spike; empty body `{}` returns 400 | | POST | /simulations/{simulationId}/inject-failure | x402 \| Bearer (write) | Fail a specific node — supply `resourceId` (preferred) or `resourceName` for deterministic targeting; pass `{"random": true}` to opt into random node selection (non-deterministic); empty body `{}` returns 400. Quick injection parks the compute/Kubernetes node for positive `recoveryPolicy.failureParkSteps` steps (falling back to `criticalSteps`), keeps `routedRps: 0`, `isRoutable: false`, and `availabilityState: "unavailable"` during that window, then allows an organic rejoin and cooldown. Explicit zero is rejected. This is not persistent `instance_down`: use `/failures` with `instance_down` plus `recover-resource` when the node must stay unavailable until an explicit action; `instance_kill` is permanent. Poll `recoveryProgress`, `availabilityState`, `isRoutable`, and `routedRps` after either flow. Response echoes `resolvedResourceId`, `resolvedResourceName`, `previousHealth`, and a `healthyNodes` array listing every currently healthy resource `{id, name}`. **Agents: call `GET /simulations/{simulationId}` first and pick a `resourceName` or `resourceId` from the `resources` array before calling this endpoint to avoid guessing invalid names.** | | GET | /simulations/{simulationId}/patterns | Bearer (read) | List traffic patterns (ramp, burst, step, wave, custom) for the simulation | | POST | /simulations/{simulationId}/patterns | Bearer (write) | Create a new traffic pattern | | PATCH | /patterns/{patternId} | Bearer (write) | Update a traffic pattern | | DELETE | /patterns/{patternId} | Bearer (write) | Delete a traffic pattern | | GET | /simulations/{simulationId}/failures | Bearer (read) | List scheduled failure injections | | POST | /simulations/{simulationId}/failures | Bearer (write) or x402 (`simulation.inject_failure_create`, $0.0010) | Schedule a failure injection (instance_kill, instance_down, az_outage, database_overload, network_latency, spot_interruption) | | PATCH | /failures/{failureId} | Bearer (write) or x402 (`simulation.inject_failure_update`, $0.0010) | Update a failure injection (e.g. deactivate early) | | DELETE | /failures/{failureId} | Bearer (write) or x402 (`simulation.inject_failure_delete`, $0.0010) | Delete a failure injection | **POST /simulations/{simulationId}/failures — request body** Create accepts **either a single failure object or a JSON array of failure objects** in one request. Sending an array creates all failure injections at once (handy for setting up a multi-fault chaos scenario in a single call). The response mirrors the request: a single object returns the created failure object (`201`); an array returns an array of created failure objects (`201`), in the same order. Validation is **all-or-nothing**: every element is validated before any is persisted, so if any element is invalid the whole request fails with a `400` and nothing is created. For an array, the offending element's JSON-pointer is prefixed with its index and the message is prefixed with `failures[]: `. For example, if the element at index `2` is missing its required `name`, the error is: ```json { "error": { "code": "MISSING_REQUIRED", "pointer": "/2/name", "message": "failures[2]: Required" } } ``` A single-object request keeps the original error shape (pointers like `/name`, no index prefix). An empty array (`[]`) returns a `400` asking for at least one failure object. ```json [ { "name": "Kill web node", "type": "instance_kill", "startTime": 10, "severity": "severe" }, { "name": "AZ outage us-east-1a", "type": "az_outage", "targetZone": "us-east-1a", "startTime": 20 }, { "name": "DB connection storm", "type": "database_overload", "startTime": 30, "parameters": { "latencyMs": 250 } } ] ``` The `type` must be one of `instance_kill | instance_down | az_outage | database_overload | network_latency | spot_interruption`. `instance_kill` PERMANENTLY removes the instance (deleting the failure does not restore it); `instance_down` is a reversible single-node outage — the node is marked critical without being removed, and deactivating (`PATCH` with `isActive: false`) or deleting the failure restores it to healthy. `spot_interruption` targets an AWS Kubernetes Spot resource and exposes a deterministic 120-second simulated notice window; configure `parameters.spotInterruption` with workload replicas, image size/cache, pull bandwidth, scheduling capacity, startup time, and `reschedule`, `drain-only`, or `fail-fast` handling. Its step metrics expose notice/deadline, pod states, pull/startup progress, unavailable/replacement nodes, and deadline misses. `severity` (`minor | moderate | severe`, default `moderate`) and `startTime` (the step the failure begins) apply to all types; `targetResourceId`, `targetZone`, `duration`, `endTime`, and the nested `parameters` object are optional. `simulationId` is injected by the server; do not include it in the body. **POST /simulations/{simulationId}/patterns — request body** Create accepts **either a single pattern object or a JSON array of pattern objects** in one request. Sending an array creates all patterns at once (handy for multi-phase profiles like ramp → burst → wave). The response mirrors the request: a single object returns the created pattern object (`201`); an array returns an array of created pattern objects (`201`), in the same order. Validation is **all-or-nothing**: every element is validated before any is persisted, so if any element is invalid the whole request fails with a `400` and nothing is created. For an array, the offending element's JSON-pointer is prefixed with its index and the message is prefixed with `patterns[]: `. For example, if the element at index `2` is a ramp missing its target, the error is: ```json { "error": { "code": "MISSING_REQUIRED", "pointer": "/2/parameters/endTraffic", "message": "patterns[2]: A \"ramp\" pattern requires a target traffic level (parameters.endTraffic). ..." } } ``` A single-object request keeps the original error shape (pointers like `/parameters/endTraffic`, no index prefix). An empty array (`[]`) returns a `400` asking for at least one pattern object. ```json [ { "name": "Morning ramp", "type": "ramp", "startTime": 0, "parameters": { "startTraffic": 100, "endTraffic": 500, "duration": 30 } }, { "name": "Lunch burst", "type": "burst", "startTime": 40, "parameters": { "peakTraffic": 5000, "duration": 20 } }, { "name": "Afternoon wave", "type": "wave", "startTime": 70, "parameters": { "baseline": 2000, "amplitude": 800, "period": 120 } } ] ``` The `type` must be one of `ramp | burst | step | wave | custom`. Legacy aliases are accepted and mapped automatically: `spike`→`burst`, `sine`→`wave`, `gradual_increase`→`ramp`. Traffic targets live under a nested `parameters` object (flat top-level fields like `endTraffic` are also accepted and folded into `parameters` for convenience). Each type validates its required fields on create and returns a clear `400` (naming the missing field, with a working example) when one is absent. The pattern moves traffic toward its target — it never adds traffic on top of the live level — so an existing simulation at 100 RPS with a ramp to 500 settles at 500, not a runaway value. ```json { "name": "Morning ramp", "type": "ramp", "startTime": 0, "parameters": { "startTraffic": 100, "endTraffic": 500, "duration": 30 } } ``` | `type` | Required `parameters` | Optional `parameters` (with defaults) | Meaning | |---|---|---|---| | `ramp` (alias `gradual_increase`) | `endTraffic` | `startTraffic` (current traffic), `duration` (60 steps; alias `durationSteps`) | Linearly move traffic to `endTraffic` over `duration` steps | | `step` | `endTraffic` | — | Immediately jump traffic to `endTraffic` | | `burst` (alias `spike`) | `peakTraffic` (alias `burstTraffic`) | `duration` (30 steps; alias `durationSteps`) | Briefly peak traffic at `peakTraffic`, then return to baseline | | `wave` (alias `sine`) | `amplitude` | `baseline` (current traffic; alias `baseTraffic`), `period` (120 steps) | Oscillate traffic around `baseline` by ±`amplitude` every `period` steps | | `custom` | `points` (non-empty array of `{ "time": , "traffic": }`) | — | Linearly interpolate traffic through the supplied points | **Observability** | Method | Path | Auth | Description | |---|---|---|---| | GET | /simulations/{simulationId}/snapshot | Bearer (read) | Single-call run object: architecture + latest metrics + active failures + last 10 significant events + _links. schemaVersion=2. Optional ?since= attaches a diff block (architecture adds/removes/changes, per-metric from/to/delta, failure adds/resolves, recommendation adds/resolves). Returns 404 if pin not found. | | GET | /simulations/{simulationId}/snapshots | Bearer (read) | List pinned snapshot summaries (id, label, pinnedAt, simulationId), newest first. Returns empty array when none pinned. | | GET | /simulations/{simulationId}/snapshots/{pinId} | Bearer (read) | Fetch the full SimulationSnapshot payload (architecture, metrics, failures, events, recommendations) for a specific pinned snapshot. 404 if pin not found. | | POST | /simulations/{simulationId}/snapshots | Bearer (write) | Pin the current simulation state; returns { pinnedSnapshotId, pinnedAt, label, simulationId }. Up to 20 pins retained per simulation. Pass pinnedSnapshotId as ?since= to GET /snapshot for a diff. | | GET | /simulations/{simulationId}/metrics | Bearer (read) | Full time-series metrics history (CPU, latency percentiles, error rate, throughput, cost/hr, connectionPressure) | | GET | /simulations/{simulationId}/events | Bearer (read) | Chronological event log (autoscaling, failures, cost spikes, manual annotations) | | POST | /simulations/{simulationId}/events | Bearer (write) | Manually inject an event annotation into the log | | DELETE | /simulations/{simulationId}/events | Bearer (write) | Clear all events from the simulation event log | **POST /simulations/{simulationId}/events — request body** ```json { "severity": "info", "message": "Deploying v2.3.1 — watching for latency regression", "resource": "web-tier" } ``` Required fields: `severity` (`"info"`, `"success"`, `"warning"`, or `"error"`) and `message` (free-text annotation string). `resource` is optional and names the affected resource. `simulationId` is injected by the server; do not include it in the body. The response (`201 Created`) returns the stored event object: `{ id, simulationId, timestamp, severity, message, resource? }`. **DELETE /simulations/{simulationId}/events** returns `204 No Content` with an empty body. All events are permanently removed from the simulation event log. **AI Analysis** | Method | Path | Auth | Description | |---|---|---|---| | POST | /simulations/{simulationId}/explain | Bearer (write) or x402 (`ai.explain`, $0.0010) | GPT-5: natural-language explanation of current simulation behaviour | | POST | /simulations/{simulationId}/optimize | Bearer (write) or x402 (`ai.optimize`, $0.0010) | GPT-5: prioritised infrastructure optimisation suggestions | | POST | /simulations/{simulationId}/troubleshoot | Bearer (write) or x402 (`ai.troubleshoot`, $0.0010) | GPT-5: step-by-step troubleshooting for a user-described issue | | POST | /simulations/{simulationId}/analyze-bottlenecks | Bearer (write) or x402 (`ai_bottleneck`, $0.0010) | GPT-5: identify performance bottlenecks and DigitalOcean migration recommendations | | POST | /simulations/{simulationId}/explain-autoscaling | Optional | GPT-5: explain why autoscaling actions were or were not taken | | POST | /ai-jobs | Bearer (write) or x402 (`ai_analysis`, $0.0010) | Submit async AI analysis job (explain/troubleshoot/analyze_bottlenecks/optimize); returns jobId immediately, poll /ai-jobs/{jobId} | | GET | /ai-jobs/{jobId} | Bearer (read) or x402 (`ai.status`, $0.0010) | Poll async AI job status: pending → running → completed/failed/cancelled | | GET | /ai-jobs/{jobId}/results | Bearer (read) or x402 (`ai.results`, $0.0010) | Retrieve results of a completed AI analysis job | | DELETE | /ai-jobs/{jobId} | Bearer (write) | Cancel a pending or running AI analysis job | **Right-Sizing & Accuracy** | Method | Path | Auth | Description | |---|---|---|---| | POST | /simulations/{simulationId}/bulk-resize | Bearer (write) or x402 (`simulation.resize`, $0.0010) | Resize all compute resources to a specified DigitalOcean Droplet size slug. DigitalOcean-only: returns 400 `PROVIDER_MISMATCH` (no mutation) when compute resources use another provider | | POST | /simulations/{simulationId}/recover-resource | Bearer (write) or x402 (`simulation.recover_resource`, $0.0010) | Recover a single failed resource by `resourceName` or `resourceId` (any provider). Clears its failure state (reversible failures only: `instance_down`, `database_overload` — `instance_kill` is permanent and returns 400 RESOURCE_KILLED); response echoes `stepsToHealthy` — a best-case estimate of the steps to poll before asserting health (assumes CPU stays below the demotion threshold on every subsequent step; if CPU exceeds the threshold the cooldown resets and more steps will be required). Other failed resources are unaffected | | POST | /simulations/{simulationId}/deploy | Bearer (write) or x402 (`simulation.deploy`, $0.0010) | Start a rolling replacement of an ECS Fargate task fleet. Target an `ecsFargate` resource by `resourceId` or `resourceName`, or omit both to auto-detect the first | | POST | /ui/simulations/{simulationId}/deploy | UI session cookie (`cwm_ui_session`) | Start the same rolling ECS Fargate deployment from a browser workspace. The cookie is issued by `POST /ui/simulations`; no API key or x402 payment is required | | GET | /simulations/{simulationId}/right-sizing-hint | Bearer (read) or x402 (`right_sizing_hint`, $0.0010) | Recommend smaller resource sizes for over-provisioned nodes (with savings %); also returns cost-only `rdsSqlServerComparisons` for bounded SQL Server Express resources, independently of `hasHint` | | POST | /simulations/{simulationId}/apply-right-sizing | Bearer (write) or x402 (`simulation.apply_right_sizing`, $0.0010) | Apply a right-sizing hint directly by `resourceId` + `recommendedSlug`; validates the slug against the resource's provider catalog before mutating — returns 400 `UNKNOWN_SIZE` with valid slug list on mismatch; cross-provider DO alternatives are accepted and store a corrected `costMultiplier` so subsequent steps bill the actual DO tier price | | GET | /simulations/{simulationId}/validate-cost-accuracy | Bearer (read) or x402 (`validate_cost_accuracy`, $0.0010) | Validate simulated cost/hr is within ±10% of provider benchmark. Agent contract: only trust `valid` when `checked: true`; `checked: false` means no benchmark reference matched and `valid: true` is a vacuous truth. | | GET | /simulations/{simulationId}/validate-performance-accuracy | Bearer (read) or x402 (`validate_performance_accuracy`, $0.0010) | Validate simulated throughput/latency is within ±15% of provider benchmark. Agent contract: only trust `valid` when `checked: true`; `checked: false` means no benchmark reference matched and `valid: true` is a vacuous truth. | | GET | /simulations/{simulationId}/validate-accuracy | Bearer (read) or x402 (`benchmark.validate`, $0.0010) | Combined cost + performance accuracy check; returns `overallValid` flag. Agent contract: only trust `overallValid` when `checked: true`; `checked: false` means nothing was compared. | **Accuracy validator response fields — `skippedCount` and `skippedReasons`** All three accuracy validators (`validate-cost-accuracy`, `validate-performance-accuracy`, `validate-accuracy`) include two fields that explain which resources were excluded from the check: | Field | Type | Description | |---|---|---| | `skippedCount` | integer | Number of resources (or resource-metric pairs) that were evaluated but had no matching benchmark entry and were therefore skipped. A non-zero value means the `valid` / `overallValid` verdict covers only the resources that were checked — not the full set. | | `skippedReasons` | string[] | Human-readable explanation for each skipped resource. Common reasons: no `serviceFamily` set, or no pricing benchmark found for the provider/serviceFamily/size combination. Example: `['"cache-layer": no serviceFamily set']`. | Agent contract: always inspect `skippedCount` alongside `checked` — `checked: true` with `skippedCount > 0` means the verdict is valid but partial (some resources were compared; others were not). `checked: false` with `skippedCount > 0` means every resource was skipped and nothing was compared. ### Scenarios API Pre-built infrastructure scenario templates. No authentication required. | Method | Path | Auth | Description | |---|---|---|---| | GET | /scenarios | Public | List all scenario templates (name, category, provider, resources) | | GET | /scenarios/{scenarioId} | Public | Get full details of a single scenario template including connections | ### Simulation Fidelity API Exposes the benchmark sources and accuracy thresholds used to calibrate the simulation engine. No authentication required. | Method | Path | Auth | Description | |---|---|---|---| | GET | /fidelity | Public | Per-provider benchmark data (pricing and performance) and accuracy thresholds (cost ±10%, performance ±15%) | | GET | /accuracy-benchmark | Public | Return the owned cwm-bench AWS benchmark for ALB → 2× m5.large → db.r5.large MySQL in us-east-2; exposes goodput/diagnostics, fit/holdout provenance, and calibrated engine comparison fields | | GET | /accuracy-benchmark/openshift | Public | Return independently sourced ROSA, ARO, OpenShift Dedicated, and self-managed OpenShift reference scenarios; platform cost is estimated and performance components are extrapolated, so no OpenShift score is claimed | | POST | /accuracy-benchmark | Public | Run the accuracy benchmark with a custom provider, compute type, and database type; supports `aws`, `gcp`, `azure`, `oci`, and `digitalocean`; see request body details below | **POST /accuracy-benchmark — request body** Required fields: `provider` (`"aws"`, `"gcp"`, `"azure"`, `"oci"`, or `"digitalocean"`), `computeCount` (integer, 1–10), `computeType` (provider-specific string), `dbType` (provider-specific string). Valid values by provider: | provider | computeType | dbType | |---|---|---| | `aws` | `t2.micro`, `m5.large`, `m5.xlarge`, `m6i.large`, `m7i.large`, `m8i.large`, `m9g.large` | `db.t3.small`, `db.r5.large` | | `gcp` | `e2-medium`, `n2-standard-4` | `db-f1-micro`, `db-n1-standard-2` | | `azure` | `Standard_B2s`, `Standard_D4s_v3` | `sql-basic`, `sql-s2` | | `oci` | `VM.Standard.E4.Flex`, `VM.Standard3.Flex` | `mysql-heatwave`, `autonomous-db-std` | | `digitalocean` | `s-2vcpu-4gb`, `c-4` | `db-s-1vcpu-1gb`, `db-s-2vcpu-4gb` | **AWS example** (canonical architecture, 2× m5.large + db.r5.large): ```json { "provider": "aws", "computeCount": 2, "computeType": "m5.large", "dbType": "db.r5.large" } ``` **GCP example** (2× n2-standard-4 Compute Engine + Cloud SQL db-n1-standard-2): ```json { "provider": "gcp", "computeCount": 2, "computeType": "n2-standard-4", "dbType": "db-n1-standard-2" } ``` **Azure example** (2× Standard_D4s_v3 VM + Azure SQL Standard S2): ```json { "provider": "azure", "computeCount": 2, "computeType": "Standard_D4s_v3", "dbType": "sql-s2" } ``` **OCI example** (2× VM.Standard3.Flex + MySQL HeatWave): ```json { "provider": "oci", "computeCount": 2, "computeType": "VM.Standard3.Flex", "dbType": "mysql-heatwave" } ``` **DigitalOcean example** (2× s-2vcpu-4gb Droplet + Managed Database db-s-2vcpu-4gb): ```json { "provider": "digitalocean", "computeCount": 2, "computeType": "s-2vcpu-4gb", "dbType": "db-s-2vcpu-4gb" } ``` The response shape is identical to `GET /accuracy-benchmark` for the scenario and provenance fields, plus `"isCustomConfig": true`. Custom requests remain scored with simulated/comparison metrics. The canonical AWS GET compares seeded predictions with owned app-host CPU and k6 in-VPC internal-load-balancer latency, owned goodput/errors, and AWS us-east-2 price-list cost ($0.4545/hour). Database CPU and the derived blend are diagnostics, not the scored app-CPU basis. The fixed seed `20240601` is used for scored results. The 1,000 RPS Burst holdout measured one failure in 1,050,142 requests (0.0000952%, raw fraction 9.52e-07); 875.12 RPS whole-run goodput includes a five-minute ramp and does not imply dropped requests. Later-day and us-west-2 holdouts have no separate predictions. **Per-provider accuracy score driver notes** The `scoreDriverNote` field is returned in every `GET /accuracy-benchmark` and `POST /accuracy-benchmark` response. For canonical AWS it explains the measured app-CPU/in-VPC-latency comparison and price-list cost. For other provider/custom results it explains the provider-specific score drivers. These notes are also shown in the comparison table. ```json [ { "provider": "aws", "scoreDriverNote": "AWS canonical comparison scores owned cwm-bench app-host CPU and k6 in-VPC internal-LB latency against seeded engine predictions, with owned goodput/errors and AWS price-list cost. Idle/Normal/Peak are fit rungs; Burst is a holdout. Later-day and second-region have no separate predictions." }, { "provider": "gcp", "scoreDriverNote": "GCP uses its own latency and error coefficients fit to e2-standard-2 / db-n1-standard-2 reference data, plus a +5 ms idle base-latency offset (Cloud LB forwarding over AWS ALB) so idle P50 reproduces the documented floor. Latency, error rate (including Burst) and low/mid-load CPU track references closely; the main remaining gap is the shared (AWS-calibrated) database CPU curve, which slightly over-predicts CPU at Burst." }, { "provider": "azure", "scoreDriverNote": "Azure uses its own latency and error coefficients fit to Standard_D2s_v3 / SQL Standard S2 reference data, plus a +3.75 ms idle base-latency offset so idle P50 reproduces the documented floor. Latency and error rate (including Burst) track references closely; the largest remaining divergence is at Burst, where the shared (AWS-calibrated) database CPU saturation curve over-predicts CPU for the 200-connection S2 tier." }, { "provider": "oci", "scoreDriverNote": "OCI uses its own latency, error, CPU, tail and throughput-shed coefficients fit to VM.Standard.E4.Flex 2-OCPU / MySQL HeatWave reference data, plus a +6.25 ms idle base-latency offset (EPYC + Flexible LB overhead) so idle P50 reproduces the documented floor. HeatWave's low in-memory error rates are captured by the fitted polynomial (Burst error matches closely), and a connection-pool goodput-shed term reproduces the documented Peak/Burst effective throughput (which sits below traffic × (1 − errorRate) due to admission back-pressure). The residual is in CPU: OCI's reference curve is unusually steep (low Normal 12 % yet high Burst 78 %), which the shared √-shaped compute curve plus the shared database CPU curve cannot fully track — leaving CPU slightly over-predicted at Normal and under-predicted at Peak/Burst." }, { "provider": "digitalocean", "scoreDriverNote": "DigitalOcean uses its own latency and error coefficients fit to s-4vcpu-8gb / db-s-2vcpu-4gb reference data, plus a +9 ms idle base-latency offset (shared-tenant network fabric) so idle P50 reproduces the documented floor. Latency and CPU track references across all tiers; the main remaining divergence is the very low documented error rates at Idle/Normal, where sub-0.5 % absolute references make small noise-floor differences a large relative delta." } ] ``` #### Fidelity notes — request-based and serverless resources Several AWS resources use a billing unit that cannot be directly expressed as a flat hourly instance rate. These resources are **excluded from the FR-12 ±10% cost-accuracy check**; some use flat approximations and Aurora Serverless v2 uses modeled ACU consumption. **Aurora Serverless v2 (`aurora-serverless`) — ACU-based cost model** Aurora Serverless v2 bills per ACU-hour ($0.12/ACU-hr in us-east-1, May 2026). The engine models each database instance's ACUs independently between `minCapacity` and `maxCapacity`, charging the existing per-ACU rate for its current modeled ACUs. Prefer those canonical input names; flat `minAcu`/`maxAcu` are accepted as compatibility aliases for the AWS Aurora Serverless v2 `db.serverless` database shape, including single- and two-instance configurations. This is a simulation of load-dependent consumption, not a measured AWS bill. For a modeled two-instance, cross-AZ Aurora Serverless v2 cluster, create an AWS database with `characteristics: { "size": "db.serverless", "serviceFamily": "aurora-serverless", "multiAz": true, "instanceCount": 2 }`. Creation materializes a separate `-reader` database resource in a different AZ with `replicaOf` pointing to the writer and sets the writer's `auroraStandbyResourceId`. Each instance has its own ACU cost row in `metrics.costBreakdown` and `metrics.serverless`; shared database storage is charged only once. A healthy reader is eligible for modeled writer promotion in quick injection and `database_crash`. This is a deterministic modeling assumption, not a real-world AWS failover-time promise. `instanceCount:1` requires `multiAz:false` and keeps one database instance and one charge. Contradictory counts, explicit duplicate replicas, and unknown AZs are rejected. Legacy `multiAz:true` without `instanceCount` does **not** establish a reader; inspect `normalizedConfig.resources[].auroraTopologyWarnings`. See https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/Concepts.AuroraHighAvailability.html and https://aws.amazon.com/rds/aurora/pricing/. Cost comparison to provisioned Aurora: | Resource | Billing model | Simulated cost | Official rate | Notes | |---|---|---|---|---| | Aurora Serverless v2 | $0.12/ACU-hr, scales with load | Modeled ACUs × rate for each active DB instance | $0.12/ACU-hr | Excluded from FR-12; consumption is simulated, not measured | | Aurora PostgreSQL (db.r6g.large) | Fixed hourly instance rate | $0.26/hr | $0.26/hr | Within FR-12 ±10% accuracy bound | At 1 ACU a single Aurora Serverless v2 instance costs roughly half the provisioned `aurora-postgresql` (db.r6g.large) instance. Compare **all billable instances**, not just the writer: a two-instance cluster at 1 ACU each is about $0.24/hr before storage and other charges. At sustained higher ACUs its per-instance charge rises. Because Aurora Serverless v2 is excluded from FR-12, the `/simulations/{id}/validate-cost-accuracy` and `/simulations/{id}/validate-accuracy` endpoints skip it when computing the overall cost accuracy verdict. Treat cost comparisons as modeled estimates, sensitive to each instance's simulated ACUs. **Aurora DSQL (`aurora-dsql`) — DPU-second + request-unit cost model** Aurora DSQL bills on two dimensions simultaneously: compute time measured in DPU-seconds ($0.000463/DPU-second in us-east-1, May 2026) and per-operation request units charged separately for reads and writes. Because the simulation has no visibility into either the instantaneous DPU count or the request-unit volume generated by a workload, neither dimension can be computed accurately from the simulation state alone. The simulated hourly cost is therefore approximated as a **flat $0.12/hr**, representing a lightly loaded reference workload. At higher DPU counts or elevated read/write throughput the actual cost will exceed this estimate. Cost comparison to provisioned Aurora: | Resource | Billing model | Simulated cost | Official rate | Notes | |---|---|---|---|---| | Aurora DSQL | DPU-second + read/write request units | $0.12/hr (flat reference) | $0.000463/DPU-second + request units | Excluded from FR-12; approximation reflects low-load baseline only | | Aurora Serverless v2 | $0.12/ACU-hr, scales with load | $0.12/hr (flat, 1 ACU reference) | $0.12/ACU-hr | Excluded from FR-12; approximation holds at ≤1 ACU sustained load | | Aurora PostgreSQL (db.r6g.large) | Fixed hourly instance rate | $0.26/hr | $0.26/hr | Within FR-12 ±10% accuracy bound | Aurora DSQL's flat $0.12/hr simulation approximation matches the Aurora Serverless v2 approximation at the 1-ACU reference point, but the two resources have fundamentally different scaling characteristics. Aurora DSQL is a distributed, serverless relational database designed for high availability across regions; its DPU consumption rises with query concurrency and data volume in ways that are decoupled from the simulation's compute-instance model. Agents comparing `aurora-dsql` and `aurora-postgresql` costs should treat the DSQL figure as a floor — a lower-bound estimate valid only for very low-throughput workloads. Because Aurora DSQL is excluded from FR-12, the `/simulations/{id}/validate-cost-accuracy` and `/simulations/{id}/validate-accuracy` endpoints skip it when computing the overall cost accuracy verdict. Agents building cost-optimization strategies that include Aurora DSQL alongside `aurora-serverless` and `aurora-postgresql` should account for the fact that all three resources carry different uncertainty levels: `aurora-postgresql` is within ±10%, `aurora-serverless` is a lower-bound ACU estimate, and `aurora-dsql` is a lower-bound DPU+request estimate. ### Pricing History API Historical and current cloud provider pricing data. No authentication required. | Method | Path | Auth | Description | |---|---|---|---| | GET | /pricing-history | Public | Pricing history document: snapshots, recent changes, per-provider trends, and resource ID list | | GET | /pricing-history/compare | Public | Compare pricing changes between two snapshots; requires query params `from` and `to` (snapshot dates from GET /pricing-history); returns changed entries, from/to labels | | GET | /pricing-history/trend/{resourceId} | Public | Full price time-series for a specific resource type (use GET /pricing-history to discover valid resource IDs) | | GET | /admin/pricing-check-status | Public | Latest automated pricing-drift check status: maps each check identifier (e.g. aws-ec2-m5large) to pass/drift/error, with the timestamp of the last run | ### RL Training API Train reinforcement learning agents on the autoscaling simulator. Supports episode lifecycle management, multi-objective reward functions (cost, latency, availability), and provider-specific autoscaling behavior (warm-up latency, cooldown periods, scaling dynamics). | Method | Path | Auth | Description | |---|---|---|---| | POST | /rl/environments | Bearer (write) or x402 (`rl_env_create`, $0.0010) | Create a new RL environment linked to a simulation | | GET | /rl/environments/{environmentId} | Bearer (read) or x402 (`rl_env_get`, $0.0010) | Get environment state and current configuration | | GET | /rl/environments/{environmentId}/observation | Bearer (read) or x402 (`rl_env_observation`, $0.0010) | Get the latest observation; returns `{obs, metrics, resources}` where obs has decision features (rps, cpu_util, instances, traffic, currentTime, tick_seconds) and metrics has evaluation outputs (cost_usd_hr, latency_p95, error_rate, uptime, sla_violations, connection_pressure?, serverless?) | | POST | /rl/environments/{environmentId}/batch-step | Bearer / x402 | Execute up to 30 actions in one HTTP round-trip; each action counts as 1 call against the 5 000 req/hr RL quota; stops early if episode ends; returns `{results: StepResponse[]}`. On 429 read the `Retry-After` response header (seconds until reset) and back off before retrying. | | POST | /rl/environments/{environmentId}/eval-episodes | Bearer / x402 | Submit an eval-episode job (async by default). Returns `202 Accepted` immediately with `{jobId, status, createdAt}`; poll `GET /rl/environments/{environmentId}/eval-episodes/{jobId}` until `status` is `completed` or `failed`. Use `?sync=true` to run inline and get the full result directly with `200 OK`. The result shape is `{episodes, meanEvalReward, trainingTotalReward, collapseThreshold, reward_collapse}` where `reward_collapse` flags policies that exploit unpriced cost dimensions. | | GET | /rl/environments/{environmentId}/eval-episodes/{jobId} | Bearer (read) or x402 (`rl_eval_status`, $0.0010) | Retrieve the result of an async eval-episode job. Returns `202` with `{jobId, status, createdAt}` while pending or running; returns `200` with the full result merged with job metadata when `status` is `completed`; returns `200` with `{jobId, status, error, createdAt, completedAt}` when `status` is `failed`. | | POST | /rl/environments/{environmentId}/reset | Bearer (write) or x402 (`rl_env_reset`, $0.0010) | Reset environment to initial state for a new episode | | DELETE | /rl/environments/{environmentId} | Bearer (write) or x402 (`rl_env_delete`, $0.0010) | Cancel and delete an RL environment; returns `{id, isActive, cancelledAt, message}` | | POST | /rl/environments/{environmentId}/step | Bearer (write) or x402 (`rl.step`, $0.0010) | Execute an action; returns `{t, obs, metrics, reward (scalar), reward_components, done, info}` — obs and metrics use same shape as the observation endpoint | | POST | /rl/environments/{environmentId}/validate-action | Bearer (read) | Pre-validate an action against budget, SLA, region, and compliance policy domains without mutating state. Returns `{allowed, outcome, mode, violations, evaluatedAt}`. Spike endpoint — no step is incremented, no events are persisted. | **Observation space (`obs`)**: rps (req/s), cpu_util (0-1 fraction), instances (compute count), traffic, currentTime, tick_seconds (simulation seconds per step), unmodeled_dimensions_active (boolean — true when evalCostOverrides are active; signals an eval episode to the policy), warmup_factor (optional 0-1 capacity factor — 1.0 when all compute instances are fully warmed; below 1.0 while any instance is still in its JIT warm-up window; absent when no resources carry a warm-up counter), routedRps_per_node (optional number — instance-count-weighted average requests per second routed to each active compute node this step; absent on reset observations and when no compute resources exist; use this instead of rps/instances to get accurate per-node load after partial failures redistribute traffic), unavailable_node_count (integer — count of compute/kubernetes/database resources whose availabilityState is "unavailable" at this step; zero means all tracked resources are serving normally; use as a compact failure-detection signal — when non-zero at least one node has failed and survivors absorb its redirected traffic; for per-node detail inspect observation.resources[].availabilityState). **Evaluation metrics (`metrics`)**: cost_usd_hr, latency_p95 (ms), error_rate (0-1), uptime (0-1), sla_violations, connection_pressure (DB pool ratio). The connection_pressure metric is optional and only present when the simulation contains database resources — it is the activeConnections / maxConnections ratio capped at 3.0; values above 1.0 indicate pool exhaustion; use it to build DB-aware reward functions. The optional `serverless` array gives a **per-DB breakdown** for each Aurora Serverless v2 database in the simulation (one entry per serverless DB; omitted entirely when none exist) — each entry has `resourceId`, `name`, `cpu_util` (0-1 fraction), `acu`, `min_acu`, `max_acu`, `warming` (true during the ~4-step ACU warm-up window when this DB's cpu_util is 0), `cost_usd_hr` (this DB's own ACU-proportional cost, ~$0.06/hr at the 0.5-ACU floor), and `connection_pressure` (this DB's own pool ratio against its fixed connection limit, either supplied as maxConnections or estimated from maximum configured ACU). Use it because the sim-wide `cpu_util` and `cost_usd_hr` aggregates average across all resources and mask a single serverless DB's cpu=0 warming window and ACU-floor cost, and the sim-wide `connection_pressure` dilutes one DB's pool saturation across every database. The optional `databases` array gives a **per-DB summary** for every database resource in the simulation (one entry per database; omitted when none exist) — each entry has `resource_id`, `name`, `cpu_util` (0-1 fraction), `max_connections`, `connection_pressure` (pool ratio), and `is_serverless` (boolean). Unlike `serverless`, which is Aurora Serverless v2-only and includes ACU fields, `databases` covers all database types including provisioned RDS, Cloud SQL, and Azure SQL. Use `databases` when you need a uniform per-DB view across a mixed-engine topology. **Per-resource recovery policy in observations**: Every resource object in the `resources` array of the observation (returned by both the observation endpoint and step responses) includes a `recoveryPolicy` field with four fields: `criticalCpuThreshold` (default 80), `criticalSteps` (default 4), `warningCpuThreshold` (default 70), `warningSteps` (default 3). Resources that have never had `set_recovery_policy` applied will carry the global defaults. Agents can read `observation.resources[i].recoveryPolicy` to compare healing configurations across resources, confirm a `set_recovery_policy` action took effect, or detect policy drift between resources serving similar roles. **Action types**: `scale_out`, `scale_in`, `adjust_threshold`, `add_resource`, `remove_resource`, `no_op`, `set_recovery_policy`. | Action | Parameters | Description | |---|---|---| | `scale_out` | `instanceCount` (int) or `instances` (int alias) | Add compute instances to increase capacity | | `scale_in` | `instanceCount` (int) or `instances` (int alias) | Remove compute instances to reduce cost | | `adjust_threshold` | `cpuThreshold` (0–100) | Change the autoscaling CPU trigger threshold | | `add_resource` | `resourceType`, `provider`, `serviceFamily`, `config` | Provision a new resource (compute or database) into the running simulation | | `remove_resource` | `resourceId` (string) | Remove an existing resource from the simulation | | `no_op` | *(none)* | Advance the simulation clock by one tick (or `tick_seconds`) without triggering any resource change. Recommended during the aurora-serverless ACU ramp window (steps 1–4) when `obs.cpu_util` reads `0` and no autoscaling decision is warranted. | | `set_recovery_policy` | `resourceId` (string), `recoveryPolicy` (object) | Override the recovery thresholds for a specific resource. `recoveryPolicy` fields: `criticalCpuThreshold` (0–100, default 80), `criticalSteps` (int ≥ 1, default 4), `warningCpuThreshold` (0–100, default 70), `warningSteps` (int ≥ 1, default 3). Use this to vary how quickly individual resources heal — stateless workloads benefit from aggressive (low step count) recovery; stateful workloads may need conservative (high step count) thresholds to avoid flapping. **Policies set via this action are persisted in the environment row and automatically re-applied on every subsequent `/reset` call — agents do not need to re-apply them after each episode.** | **Action parameter note — `instanceCount` vs `obs.instances`**: The `instanceCount` parameter in `scale_out` and `scale_in` actions is the number of instances to add or remove (delta, not a target). The `obs.instances` field in every observation is a read-only state count — the total number of active compute instances at that step — and is not an action parameter. **Per-resource instance bounds**: A compute resource's `characteristics.maxInstances` and `characteristics.minInstances` act as hard per-resource ceilings/floors for `scale_out` and `scale_in` actions. The effective maximum is `min(autoscalingConfig.maxInstances, characteristics.maxInstances)` and the effective minimum is `max(autoscalingConfig.minInstances, characteristics.minInstances)`. When an action is trimmed to honour these bounds, the step response `info` object gains four flat keys: `scale_clamped: true` (boolean), `requested` (integer — the `instanceCount` delta the agent requested), `actual` (integer — the delta actually applied; `0` if the action was fully blocked because the fleet is already at the limit), and `limit` (integer — the effective bound that triggered clamping). These keys are absent when no clamping occurred. **RL traffic source**: Traffic in RL episodes is driven by the environment's `episodeConfig`, not by simulation `/patterns` endpoints. Patterns created on the underlying simulation (via `POST /simulations/{id}/patterns`) are not applied to RL episodes — they only affect the workspace UI simulation. To drive high-RPS load in an episode, set `episodeConfig.initialTraffic` (e.g. `"initialTraffic": 50000`) when creating or resetting the environment, or use `add_resource` with a traffic-intensive configuration. The `episodeConfig.targetTrafficPattern` field (`ramp`, `burst`, `step`, `wave`, `custom`) controls how traffic evolves during the episode independently of any sim-level patterns. **`add_resource` action — AWS database `serviceFamily` values**: Use these in the `resource.serviceFamily` field when the action type is `add_resource` and `resource.type` is `"database"` with `resource.provider` `"aws"`: | serviceFamily | Description | |---|---| | `aurora-postgresql` | Provisioned Aurora PostgreSQL cluster — fixed instance size, predictable cost | | `aurora-dsql` | Aurora DSQL — distributed serverless SQL optimized for high-concurrency, multi-region writes | | `aurora-serverless` | Aurora Serverless v2 — scales ACUs continuously between `minCapacity` and `maxCapacity`; cost tracks actual load rather than peak provisioned size; best choice when workload is variable or unpredictable | **Cost is auto-derived from shape — you do not need to send `costMultiplier`**: For provisioned AWS databases such as `{ "serviceFamily": "aurora-postgresql", "size": "db.r6g.large" }` (≈ $0.26/hr) and `{ "serviceFamily": "rds", "size": "db.r5.large" }` (≈ $0.24/hr), the engine derives the hourly rate from `serviceFamily` + `size` without `costMultiplier`. Advanced/what-if overrides apply to those fixed-rate shapes. For `{ "serviceFamily": "aurora-serverless", "size": "db.serverless", "minAcu": 2, "maxAcu": 16 }`, Aurora Standard Serverless v2 bills the modeled current ACU per instance × $0.12/ACU-hour in us-east-1, including the warm-up floor; a non-default cost multiplier or unverified region is rejected. Source: https://aws.amazon.com/rds/aurora/pricing/ (Aurora Standard us-east-1 example; verified 2026-09-25). Variable-ACU Aurora remains excluded from the fixed-instance FR-12 benchmark. **RL quickstart — Aurora Serverless v2 as a cost-optimizing `add_resource` action**: ```json { "action": "add_resource", "resource": { "name": "aurora-sv2", "type": "database", "provider": "aws", "serviceFamily": "aurora-serverless", "config": { "minCapacity": 0.5, "maxCapacity": 16 } } } ``` Use this action when the agent detects high `cost_usd_hr` and low sustained `cpu_util`, or when it wants to trial a database tier that scales to near-zero ACUs during idle periods. `minCapacity` (ACUs) sets the floor; `maxCapacity` sets the ceiling. Reward signal will reflect the lower per-step cost once the simulated workload drops below the provisioned-instance breakeven point. **Aurora Serverless v2 — expected observation shape after `add_resource`**: The simulation models a **2–4 step ACU scale-out delay** before the resource reaches operating capacity. During this ramp period `cpu_util` for the new resource reads `0` and `cost_usd_hr` reflects only the `minCapacity` ACU floor (≈ $0.06/hr at 0.5 ACUs in us-east-1). Once the simulated workload is routed to the resource, ACUs scale proportionally to load; the engine advances ACUs in discrete increments per step, so `cost_usd_hr` rises incrementally over **3–6 steps** rather than jumping immediately to the steady-state value. At steady state, `cpu_util` stabilises in the **20–70 %** range for typical OLTP workloads, `latency_ms` falls relative to a saturated provisioned Aurora instance, and `cost_usd_hr` settles between the floor cost and the full `maxCapacity` rate (≈ $1.92/hr at 16 ACUs). When load drops below the `minCapacity` breakeven, `cpu_util` returns to near `0` and `cost_usd_hr` reverts to the floor within **1–2 steps**. Tune reward functions to account for the ramp lag: penalise latency violations during the first 4 steps after the action rather than treating them as a steady-state failure. **Aurora Serverless v2 — episode-loop quickstart (Python)**: ```python import requests BASE = "https://www.cloudworldmodel.ai/api" HEADERS = {"Authorization": "Bearer ", "Content-Type": "application/json"} # 1. Create an RL environment linked to an existing simulation that already contains # at least one AWS database resource (the engine clones it for the new step). env = requests.post(f"{BASE}/rl/environments", json={"simulationId": "", "episodeConfig": {"maxSteps": 40, "tick_seconds": 60}}, headers=HEADERS).json() env_id = env["environment"]["id"] # 2. add_resource action — adds an AWS database resource to the running simulation. # The engine models a 2–4 step ACU scale-out delay for aurora-serverless targets. # Step request shape: {"action": {"type": "...", "parameters": {...}}, "tick_seconds": N} step1 = requests.post(f"{BASE}/rl/environments/{env_id}/step", headers=HEADERS, json={ "action": { "type": "add_resource", "parameters": {"resourceType": "database", "provider": "aws"} }, "tick_seconds": 60 }).json() # Step response keys: t, obs, metrics, reward, reward_components, done t = step1["t"] # step index (1 after first action) obs = step1["obs"] # {rps, cpu_util, instances, traffic, currentTime, tick_seconds} metrics = step1["metrics"] # {cost_usd_hr, latency_p95, error_rate, uptime, sla_violations} # 3. Ramp window (t = 1–4): obs.cpu_util == 0, metrics.cost_usd_hr at minCapacity floor (~$0.06/hr). # Use no_op to advance the clock without making any autoscaling decision (see Action types table). # Only count cost + stability reward components — skip performance and sla during ACU ramp-up. RAMP_STEPS = 4 result = step1 for _ in range(RAMP_STEPS): result = requests.post(f"{BASE}/rl/environments/{env_id}/step", headers=HEADERS, json={ "action": {"type": "no_op", "parameters": {}}, "tick_seconds": 60 }).json() t = result["t"] obs = result["obs"] metrics = result["metrics"] rc = result["reward_components"] # Exclude performance and sla components during ACU ramp to avoid false latency penalties adjusted_reward = rc["cost"] + rc["stability"] print(f"Ramp t={t:2d} cpu_util={obs['cpu_util']:.3f} " f"cost_usd_hr={metrics['cost_usd_hr']:.4f} adj_reward={adjusted_reward:.3f}") # 4. Steady-state training (t >= 5): all reward components are now meaningful. With a database # resource present this is five components — performance, cost, stability, sla, and the # connection_pressure DB-pool penalty (see the worked connection-pool example below). # metrics.cost_usd_hr climbs over 3–6 steps toward the ACU-proportional steady-state rate. done = result["done"] while not done: result = requests.post(f"{BASE}/rl/environments/{env_id}/step", headers=HEADERS, json={ "action": {"type": "adjust_threshold", "parameters": {"cpuThreshold": 70}}, "tick_seconds": 60 }).json() t = result["t"] obs = result["obs"] metrics = result["metrics"] reward = result["reward"] done = result["done"] print(f"Train t={t:2d} cpu_util={obs['cpu_util']:.3f} " f"latency_p95={metrics['latency_p95']:.0f}ms " f"cost_usd_hr={metrics['cost_usd_hr']:.4f} reward={reward:.3f}") # 5. Reset for the next episode requests.post(f"{BASE}/rl/environments/{env_id}/reset", headers=HEADERS) ``` Key points: - **Step request shape**: `{ "action": { "type": "", "parameters": {...} }, "tick_seconds": 60 }`. - **Step response keys**: `t` (step index), `obs` (rps, cpu_util, instances, traffic, currentTime, tick_seconds), `metrics` (cost_usd_hr, latency_p95, error_rate, uptime, sla_violations, connection_pressure?, serverless?), `reward`, `reward_components` (performance, cost, stability, sla, connection_pressure?), `done`. - **Ramp window (t = 1–4)**: `obs.cpu_util` is `0`; `metrics.cost_usd_hr` holds at the `minCapacity` floor (~$0.06/hr at 0.5 ACUs). Exclude `reward_components.performance` and `reward_components.sla` to avoid penalising latency noise during ACU scale-out. - **Steady state (t ≥ 5)**: all reward components are meaningful — performance, cost, stability, sla, plus `connection_pressure` when the simulation contains a database resource (five signals in that case); `metrics.cost_usd_hr` settles between the floor and the full `maxCapacity` rate (~$1.92/hr at 16 ACUs) proportional to actual OLTP load. - Use `tick_seconds=60` for fine-grained training; use `tick_seconds=300` for warm-up fast-forward (see tick-rate notes below). **Reward components (`reward_components`)**: performance (latency adherence), cost (cost efficiency), stability (autoscaling stability penalty), sla (SLA violation penalty), unmodeled_cost (extra penalty from evalCostOverrides — egress/cross_az_traffic surcharges applied only inside eval episodes; absent in normal training; reveals blind spots when eval reward drops vs training reward), connection_pressure (DB pool saturation penalty — only present when database resources exist; 0 when pool pressure ≤ 1.0, decreases with slope −1 per unit from 1.0 to 1.5 reaching −0.5, then drops with slope −2 per unit above 1.5 — twice as steep — flooring at −1.0 at pressure ≥ 1.75; added directly to the weighted sum so agents learn to scale databases before pool exhaustion. **Important — multi-DB topologies:** `reward_components.connection_pressure` uses the **worst-case (highest) per-DB connection pressure** across all databases, not the sim-wide aggregate. In a primary + replica topology the aggregate dilutes the primary's saturation (e.g. a primary at CP 3.0 with a replica at CP 0.25 gives an aggregate of ~0.476, below the 1.0 threshold). Using the max ensures the agent receives a meaningful penalty even when a healthy replica masks a saturated primary. `metrics.connection_pressure` in the `metrics` block is still the sim-wide aggregate — use `metrics.databases[].connection_pressure` for the full per-DB breakdown. This component is absent until a database resource first exists; when you add a DB mid-episode via add_resource, the penalty is ramped in linearly over 5 steps — starting at 1/5 of its full magnitude on the first step it appears and reaching full strength after 5 steps — so the introduction does not cause a sudden step-to-step reward discontinuity that would look like noise to your agent. Treat the first appearance as a near-zero baseline). **Reward weights (`rewardWeights`)**: Set at create or reset time to control how each reward component contributes to the scalar total. Valid create-time keys: `cost` (cost optimization), `resilience` (low error rates and SLA compliance), `latency` (low P95 latency), `performance` (throughput and capacity), `stability` (autoscaling stability, penalizes thrashing). These are distinct from the step-response `reward_components` fields (`connection_pressure`, `unmodeled_cost`, `sla`) which are computed automatically and not user-weighted. Values are normalized server-side to sum to 1.0 — send raw priority scores or pre-normalized fractions. Example: `{ "cost": 0.4, "resilience": 0.2, "latency": 0.2, "performance": 0.1, "stability": 0.1 }`. When omitted the built-in default weighting applies (latency 0.35 / resilience 0.40 / cost 0.25). **batch-step vs eval-episodes request body**: These two multi-action endpoints use different body schemas. `POST .../batch-step` uses a `steps` array — each element is `{"action": {...}, "tick_seconds": N}` — and executes actions sequentially in the live environment. `POST .../eval-episodes` uses an `actions` 2-D array — each outer element is one episode (an ordered list of `Action` objects) — and runs ephemeral rollouts without mutating the stored environment state. Example eval-episodes request body: `{"actions": [[{"type": "no_op", "parameters": {}}, {"type": "scale_out", "parameters": {"instanceCount": 1}}]], "collapseThreshold": 0.20}`. **Worked example — the connection-pool penalty when a database is overloaded**: When the simulation contains a database resource, every step also reports `metrics.connection_pressure` (the `activeConnections / maxConnections` ratio, capped at 3.0) and adds the fifth `reward_components.connection_pressure` signal. As traffic climbs past what the database's pool can absorb, the ratio crosses 1.0 and the penalty turns negative. For example, after the agent lets RPS rise without scaling the database, a step might return: ```json { "t": 12, "obs": { "rps": 1800, "cpu_util": 0.74, "instances": 3, "traffic": 1800, "currentTime": 720, "tick_seconds": 60 }, "metrics": { "cost_usd_hr": 0.42, "latency_p95": 240, "error_rate": 0.02, "uptime": 0.998, "sla_violations": 0, "connection_pressure": 1.3 }, "reward": -0.18, "reward_components": { "performance": 0.12, "cost": -0.05, "stability": 0.05, "sla": 0.0, "connection_pressure": -0.3 }, "done": false } ``` Here `metrics.connection_pressure` is `1.3` (the pool is at 130% of capacity), so `reward_components.connection_pressure` is `-0.3` — on the slope −1 segment, `-(1.3 − 1.0) = −0.3`. That penalty is added straight onto the weighted base reward, dragging the total `reward` negative even though `performance` is still positive. The lesson for an agent: when `metrics.connection_pressure` approaches 1.0, act on the **database** (scale it up for a higher `maxConnections`, add read replicas to spread connections, add a cache to cut query volume, or use a connection pooler) rather than only scaling the compute tier — scaling web instances alone does not relieve pool saturation and the `connection_pressure` penalty will keep growing (steeper above 1.5, flooring at −1.0 by pressure ≥ 1.75). **Worked example — the per-DB `serverless` breakdown in a multi-resource simulation**: When the simulation mixes a busy compute tier with an Aurora Serverless v2 database, the sim-wide `metrics.cpu_util` and `metrics.cost_usd_hr` are aggregates that average across every resource, so a single serverless DB sitting in its cpu=0 warm-up window or pinned at the ~$0.06/hr ACU floor is invisible at the top level, and the sim-wide `connection_pressure` spreads one DB's pool saturation across all databases. The optional `metrics.serverless` array exposes each serverless DB on its own. For example, while the compute tier is saturated during the warm-up window a step might return: ```json { "t": 3, "obs": { "rps": 5000, "cpu_util": 0.95, "instances": 2, "traffic": 5000, "currentTime": 180, "tick_seconds": 60 }, "metrics": { "cost_usd_hr": 0.21, "latency_p95": 180, "error_rate": 0.01, "uptime": 0.999, "sla_violations": 0, "connection_pressure": 0.8, "serverless": [ { "resource_id": "aurora-sl", "name": "Aurora Serverless DB", "cpu_util": 0.0, "acu": 0.5, "min_acu": 0.5, "max_acu": 16, "warming": true, "cost_usd_hr": 0.06, "connection_pressure": 1.4 } ] }, "reward": -0.12, "done": false } ``` Even though sim-wide `cpu_util` is dominated by the compute tier, the `serverless[0]` entry reveals whether the DB is still warming (`cpu_util` 0) and parked at the `cost_usd_hr` floor. Its `connection_pressure` uses a fixed limit while ACU scales. An agent training on multi-resource topologies should read `metrics.serverless[i]` for per-DB warming, cost, and pool signals rather than relying on masked aggregates. **`metrics.connection_pressure` in multi-DB topologies**: The sim-wide `metrics.connection_pressure` is computed as total estimated connections divided by the **sum** of every database's pool. In a single-DB simulation it equals that database's own saturation ratio directly. When the simulation contains an Aurora Serverless v2 database, use the `metrics.serverless` array to inspect per-DB `connection_pressure`, `acu`, and `warming` state alongside the sim-wide aggregate — the `serverless` array is the right per-DB diagnostic tool, not a separate `databases` field. **Per-step tick rate (`tick_seconds`)**: Each step request may include an optional `tick_seconds` integer (1–3600) to override the episode-level clock advancement for that step only. Use large values (e.g. 300 s) during a warm-up phase to fast-forward through startup noise, then switch to smaller values (e.g. 60 s) for fine-grained training. The actual value used is reflected in `observation.tick_seconds` in the response. Example two-phase pattern: run 20 warm-up steps at `tick_seconds=300` (covering ~1.7 h of simulated time), then switch to `tick_seconds=60` for the main training loop. **Provider-specific notes**: - AWS: EC2 warm-up ~30 s; default cooldown 300 s. - AWS (`aurora-serverless`): 2–4 step ACU scale-out delay after `add_resource`; `cost_usd_hr` ramps over 3–6 steps to steady state; `cpu_util` 20–70 % at steady state for OLTP; reverts to `minCapacity` floor cost within 1–2 steps of idle. See "Aurora Serverless v2 — expected observation shape" above for full details. - GCP: Managed instance group warm-up ~45 s; default cooldown 120 s. - Azure: VM Scale Set provisioning ~60 s; default cooldown 300 s. - OCI: Flex shape scaling modeled with OCPU granularity; Autonomous DB burst handled automatically. - DigitalOcean: Droplet cold-start ~30 s; default cooldown 180 s. ### Infrastructure Optimization API Analyzes a simulation architecture and generates tested variations ranked by cost and performance. | Method | Path | Auth | Description | |---|---|---|---| | POST | /analysis/optimize | Bearer (write) or x402 (`optimization.run`, $0.0050) | Submit an infrastructure optimization job | | GET | /analysis/jobs/{id} | Bearer | Get optimization job status | | GET | /analysis/jobs/{id}/recommendations | Bearer or x402 (`ai.recommendations`, $0.0010) | Get ranked optimization recommendations | ### Predictive Scaling API Validate infrastructure against traffic forecasts and optimize autoscaling thresholds. | Method | Path | Auth | Description | |---|---|---|---| | POST | /predictions/validate | Bearer (write) or x402 (`prediction.validate`, $0.0050) | Submit a predictive scaling validation job with a traffic forecast | | POST | /predictions/optimize-thresholds | Bearer (write) or x402 (`prediction.optimize_thresholds`, $0.0050) | Submit a threshold optimization job | | GET | /predictions/jobs/{jobId} | Bearer | Get prediction job status (supports polling) | | GET | /predictions/jobs/{jobId}/stream | Bearer | Subscribe to job progress via Server-Sent Events (SSE) | | GET | /predictions/jobs/{jobId}/results | Bearer | Get final prediction or optimization results | **Result payload includes**: Recommended scale-out/scale-in thresholds, projected cost at recommended settings, SLA violation risk forecast, detected bottlenecks, and confidence intervals. **Traffic forecast requirements**: At least 5 distinct load-level steps covering a minimum of 60 simulation steps, with strictly increasing timestamps and a clear ramp-up, sustained peak, and ramp-down phase. **Predictive scaling quickstart — Aurora Serverless v2 as a database resource**: ```json { "simulationId": "sim-abc123", "trafficForecast": { "name": "Aurora Serverless Peak Test", "description": "Ramp-up to peak, sustained load, ramp-down", "dataPoints": [ { "timestamp": 0, "rps": 50, "label": "Baseline" }, { "timestamp": 25, "rps": 300, "label": "Ramp-up" }, { "timestamp": 50, "rps": 800, "label": "Peak" }, { "timestamp": 75, "rps": 400, "label": "Decline" }, { "timestamp": 100, "rps": 80, "label": "Post-peak baseline" } ] }, "testSteps": 100 } ``` Include `aurora-serverless` in prediction payloads whenever the simulation contains a variable-capacity database. The engine models a **2–4 step ACU ramp delay** after a load increase before the new capacity is fully available, and a **3–6 step incremental cost delta** as ACUs increase one tier at a time toward the new steady state — meaning projected cost during a traffic ramp is lower in early steps and reaches the provisioned-equivalent level only after the ramp completes. Both effects appear in the result payload's `bottlenecks` and `costProjection` fields so agents can account for the lag when setting scale-out lead time. ### Chaos Engineering API Inject failures to discover architectural weak points. Returns resilience scores, vulnerability analysis, and remediation recommendations. | Method | Path | Auth | Description | |---|---|---|---| | GET | /chaos/scenarios | Public | List pre-built chaos scenarios | | POST | /chaos/run | Bearer (write) or x402 (`chaos.run`, $0.0050) | Submit a chaos engineering test (single failure injection) | | GET | /chaos/jobs/{jobId} | Bearer | Get chaos job status | | GET | /chaos/jobs/{jobId}/stream | Bearer | Subscribe to job progress via Server-Sent Events (SSE) | | GET | /chaos/jobs/{jobId}/results | Bearer | Get chaos results and resilience report | | POST | /chaos/batch | Bearer (write) or x402 (`chaos_batch`, $0.0050) | Submit multiple chaos jobs in parallel | | GET | /chaos/batch/{batchId} | Bearer | Get batch job status | | GET | /chaos/batch/{batchId}/results | Bearer | Get all batch results aggregated | **Supported failure types** (`customInjections[].failureType`): `kill_instance`, `zone_outage`, `database_crash`, `database_slowdown`, `database_overload`, `network_partition`, `cpu_stress`, `memory_pressure`. **Zone-failure naming — three surfaces, three names:** | Name | Where it appears | Context | |---|---|---| | `zone_outage` | `customInjections[].failureType` in `POST /chaos/run` | Async chaos API custom injection | | `az_outage` | `type` in `POST /simulations/{id}/failures` | Synchronous failure-injection API | | `zone_failure` | `scenarioId` in `POST /chaos/run` or `POST /chaos/batch` | Pre-built scenario ID | `database_overload` models **connection-pool saturation** (not a hard crash like `database_crash`): the targeted database's connection headroom shrinks and latency climbs, but the instance stays reachable. It is accepted by **both** the asynchronous chaos API (`POST /chaos/run`, as a `customInjections[].failureType` or via the prebuilt `scenarioId: "database_overload"`) **and** the synchronous failure-injection API (`POST /simulations/{id}/failures`, as a `type`). In a multi-cloud simulation, a one-sided `database_overload` only pressures the targeted database — connection load is charged on the share of traffic actually routed to each database, so a quiet, non-targeted database no longer emits spurious connection-pressure events. Aurora Serverless v2 (`aurora-serverless`) is a valid target for `database_crash` and `database_slowdown` failures. Use the resource `id` assigned when the resource was added to the simulation. For async `database_crash` custom injections (`POST /chaos/run` or each `/chaos/batch` scenario), set `promotionDelaySeconds` and/or `restartDelaySeconds` inside the injection. Alternatively, use prebuilt `scenarioId: "database_crash"` and set these fields at the top level of `/chaos/run`, or on that scenario object in `/chaos/batch`; other scenario IDs reject them. Each override applies only to its job, not to the saved scenario. Values are bounded integer simulation seconds (0–86400). The defaults are 30 seconds for a healthy explicit `replicaOf` promotion and 1800 seconds after the injection duration for an unreplicated writer restart. These are test assumptions, not provider SLAs; `multiAz` and `instanceCount` alone do not create a replica. Read effective per-target delays and `recoveryMode` in `resilienceScore.metrics.databaseCrashAssumptions` (for batch, under each child job's `resilienceScore`). **Chaos quickstart — `database_overload` via prebuilt scenario**: ```json { "simulationId": "sim-abc123", "scenarioId": "database_overload", "duration": 300 } ``` **Chaos quickstart — `database_overload` via custom injection (targeting one database)**: ```json { "simulationId": "sim-abc123", "customInjections": [ { "failureType": "database_overload", "targetResourceId": "db-primary", "intensity": 85, "duration": 300 } ] } ``` **`POST /chaos/run` response shape**: the `202` body returns a top-level `jobId` (convenience mirror of `job.id`) alongside the nested `job` object: ```json { "jobId": "job-abc123", "job": { "id": "job-abc123", "type": "chaos_test", "status": "pending", "simulationId": "sim-abc123", "createdAt": "2026-06-13T00:00:00.000Z" }, "message": "Chaos test started. Use GET /chaos/jobs/{id} to check status." } ``` **Chaos quickstart — `database_crash` targeting an Aurora Serverless v2 resource**: ```json { "simulationId": "sim-abc123", "failureType": "database_crash", "targetResourceId": "db-sv2", "durationSteps": 5 } ``` **Chaos quickstart — `database_slowdown` targeting an Aurora Serverless v2 resource**: Unlike `database_crash` (which models a hard instance failure), `database_slowdown` simulates ACU throttling and connection-pool pressure on Aurora Serverless v2 — the latency spikes and capacity exhaustion that occur when traffic exceeds the configured `maxCapacity` or when the ACU scaling lag cannot keep up with a sudden burst. Use this failure type to verify that upstream services degrade gracefully (retry budgets, circuit breakers, read-replica fallback) rather than failing outright. ```json { "simulationId": "sim-abc123", "failureType": "database_slowdown", "targetResourceId": "db-sv2", "durationSteps": 5 } ``` **Aurora Serverless v2 recovery behavior**: When `database_crash` targets an `aurora-serverless` resource, the simulation models a **2–4 step ACU warm-up period** after the crash ends before full query throughput is restored. During those warm-up steps the resource appears in the observation with reduced capacity and elevated latency even though the failure injection has technically ended. The result payload's `timeToRecover` estimate includes this warm-up window, and `affectedServices` will list upstream compute resources that remain partially degraded until ACU scaling completes. **Chaos batch quickstart — combining `database_crash` and `database_slowdown` against Aurora Serverless v2**: Submitting both failure types in a single `/chaos/batch` call lets you test two distinct failure modes in parallel rather than sequentially. `database_crash` models a hard instance failure (connection refused, immediate downstream errors), while `database_slowdown` models ACU throttling and connection-pool pressure (latency spikes, partial degradation). Running them together reveals whether your application handles both gracefully at the same time — for example, whether a circuit breaker that trips on hard failures also activates on slow responses, or whether retry logic intended for transient slowdowns masks a total outage. ```json { "simulationId": "sim-abc123", "scenarios": [ { "scenarioId": "database_crash", "duration": 300 }, { "scenarioId": "database_slowdown", "duration": 300 } ] } ``` The `202` response returns a `batchId` and a `type: "batch_chaos_test"` job object. Poll `GET /chaos/batch/{batchId}` for overall batch status and `GET /chaos/batch/{batchId}/results` for the aggregated resilience reports once all jobs complete. **Batch resilience aggregation**: When every completed child has the explicit `resilienceScore.metrics.outcome` contract, batch outcomes aggregate request totals and durations (so availability and error rate are weighted by observed or modeled offered requests), aggregate recovery state conservatively, and recompute `overall` from the aggregated, displayed breakdown components using `recovery .25 + availability .30 + dataIntegrity .25 + gracefulDegradation .20`. If any child is legacy-only, the aggregate preserves the legacy average/max aggregation and omits synthesized outcome evidence; it does not invent denominators from legacy availability percentages. If a batch mixes request-backed evidence with a child that has only the `estimated_from_error_rate` rate-only fallback, the aggregate explicitly uses `availabilityBasis: "unavailable"` and does not combine those children into an offered-request denominator. It still retains the independent goodput, request-window duration, and duration-weighted error-rate evidence. **Chaos multi-fault cascade — zone + database**: Use `customInjections` with multiple entries in a single `POST /chaos/run` call to inject correlated failures simultaneously. This is the reliable way to test whether upstream services handle both a zone outage and a database slowdown at the same time — for example, whether retry logic meant for transient slowdowns also activates correctly when the primary AZ is unreachable: ```json { "simulationId": "sim-abc123", "customInjections": [ { "failureType": "zone_outage", "targetZone": "us-east-1a", "duration": 120 }, { "failureType": "database_slowdown", "targetResourceId": "db-primary", "intensity": 80, "duration": 120 } ] } ``` **Result payload includes**: Resilience score (0–100), affected services, blast radius estimate, time-to-recover estimate, detected vulnerabilities, and remediation recommendations. #### Resilience outcome evidence Chaos results may include the backward-compatible `resilienceScore.metrics.outcome` block. Read it as the authoritative evidence when present: - `metrics.throughput` is successful throughput (goodput). `metrics.offeredRps` is aggregate original-client offered RPS for the same step/window. The normal simulation engine emits it on each final top-level step result and persists it with current metric history; it remains optional only for legacy histories. Optional `metrics.modeledShedRps` is the aggregate modeled capacity-shed RPS for that same slice and is emitted only when complete aggregate capacity evidence exists. Neither is retry, dependency-path, or per-resource RPS, and shedding is never fabricated from goodput or error rate. - `outcome.availabilityBasis` is `observed_offered_load`, `estimated_from_goodput_and_error_rate`, `mixed_observed_and_estimated`, or `estimated_from_error_rate`, `unavailable`. Availability is duration-weighted goodput divided by offered load. Thus 850 successful RPS out of 1,000 offered RPS is **85% availability** even with a 15% error rate. Missing denominators are labeled estimates or unavailable; request totals are never invented. If the denominator is reconstructed from goodput and error rate, the estimate is exposed as `outcome.modeledOfferedRequests`/`modeledOfferedLoadRps` (with deprecated `estimatedOfferedRequests`/`estimatedOfferedLoadRps` aliases), not as observed load. If reconstruction is impossible but duration-weighted error evidence exists, `estimated_from_error_rate` reports `sum((1 - errorRate / 100) × duration) / totalDuration` and leaves offered totals/rates null. With a reconstructable denominator, `errorRatePercent` is request-weighted by observed or inferred offered RPS, not a simple average of sample percentages. A 15% error rate with zero goodput and no offered load is `unavailable`, not inferred 85% availability. In a batch, mixing request-backed children with a rate-only `estimated_from_error_rate` child is also `unavailable` rather than a denominator computed from the request-backed subset; independent goodput, request-window duration, and duration-weighted error-rate evidence remain available. - `outcome.requestWindowSeconds` is the exact duration of the normalized simulation-time request window, including the final sample duration. The analyzer accepts an optional `windowEndSeconds` input as an exclusive simulation-time boundary and clips the final interval to it; the public request-window value reports the resulting duration used for weighting. - `outcome.modeledShedRequests` and `outcome.modeledSheddingRps` identify work rejected by modeled capacity. Do not count modeled shedding as application errors a second time. - Recovery requires error rate **below 5% for 30 simulation seconds**. `outcome.recovery.sustainPeriodCompleted` and nullable milestones distinguish an unhealthy unrecovered run (`sustainPeriodCompleted: false`, `finalState: "unhealthy"`, legacy recovery duration `0`) from sustained recovery (`finalState: "healthy"`). A healthy blip shorter than 30 seconds does not complete recovery; a later relapse retains the first recovery milestone, leaves `sustainPeriodCompleted: true`, clears `finalSustainedRecoveryCompletedSeconds` until another sustained recovery, and sets `regressedAfterRecovery: true`. Healthy state also requires positive goodput/service evidence, not only a low error rate. - A healthy history with no detected incident is baseline evidence, not recovery: recovery milestones remain null and `sustainPeriodCompleted` is false (the final state may still be `healthy`). - The score uses displayed rounded components at weights `recovery .25 + availability .30 + dataIntegrity .25 + gracefulDegradation .20`. A grade is not a deadline verdict or proof that an unrecovered run is production-ready. See the focused contract guide at `https://www.cloudworldmodel.ai/docs/resilience-outcome-metrics.md` and the `ResilienceScore`/`ResilienceMetrics` schemas in `https://www.cloudworldmodel.ai/openapi.yaml`. **Resilience grade caveat**: The `grade` field is derived from the live weighted `overall` score (including recovery, availability, data integrity, and graceful degradation). Vulnerability findings may be static architectural evidence and can remain similar across fault combinations, but the grade itself is not a static SPOF label and can change with observed outcomes. To inspect both, pair `POST /chaos/run` (vulnerability report) with the synchronous `POST /simulations/{id}/failures` + step-loop pattern (live metrics). ### Multi-Cloud Strategy Exploration API Generate and analyze multi-cloud architecture variants to identify optimal provider combinations. | Method | Path | Auth | Description | |---|---|---|---| | POST | /multi-cloud/explore | Bearer (write) or x402 (`multicloud.explore`, $0.0050) | Submit a multi-cloud strategy exploration job | | GET | /multi-cloud/jobs/{jobId} | Bearer (read) or x402 (`multicloud.status`, $0.0010) | Get job status (supports polling) | | GET | /multi-cloud/jobs/{jobId}/stream | Bearer (read) or x402 (`multicloud_stream`, $0.0010) | Subscribe to job progress via Server-Sent Events (SSE) | | GET | /multi-cloud/jobs/{jobId}/partial-results | Bearer (read) or x402 (`multicloud_partial_results`, $0.0010) | Get strategies accumulated so far while a job is running (isComplete: false until terminal) | | GET | /multi-cloud/jobs/{jobId}/results | Bearer (read) or x402 (`multicloud_results`, $0.0010) | Get final exploration results and comparison report. **Only returns data once the job has finished** (status `completed`). While the job is still `pending` or `running` it returns HTTP 400 `INVALID_REQUEST`; a `failed` job returns HTTP 500. To read strategies as they accumulate during a run, poll `/multi-cloud/jobs/{jobId}/partial-results` instead. | **Polling pattern**: After `POST /multi-cloud/explore`, poll `GET /multi-cloud/jobs/{jobId}` (or subscribe via `/stream`) until `status` is `completed`, then call `/results` once for the final ranked strategies and comparison report. Do not poll `/results` for in-progress data — use `/partial-results` for that. **Agent tip — request body field placement**: The request body for `POST /multi-cloud/explore` has five optional top-level fields alongside `workloadProfile`: `optimizationWeights`, `beginnerMode`, `maxCostPerHour`, `errorBudgetPct`, and `modifiers`. These are siblings of `workloadProfile`, not nested inside it. If you place any of these fields *inside* `workloadProfile`, the server silently ignores them and responds 202 — your weights and constraints will have no effect. Always send them at the top level. **`modifiers` field** (optional object): Apply pricing discounts to the exploration run. All sub-fields default to off. Sub-fields: `awsCommitment` (`"on-demand"` | `"1yr"` | `"3yr"`) — AWS Savings Plan tier (~40% off 1-yr, ~60% off 3-yr); `azureHybridBenefit` (boolean) — Azure Hybrid Benefit (~40% off Azure compute/DB for existing Microsoft licence holders); `spotEligible` (boolean) — Spot/Preemptible VM pricing (AWS ~70% off, GCP ~80% off, Azure ~75% off, OCI ~50% off; DigitalOcean has no spot offering); `oracleLicenseHolder` (boolean) — OCI BYOL note (~50% off OCI database costs for existing Oracle SE/EE licence holders, informational only). **`optimizationWeights` strict validation**: The `optimizationWeights` object rejects unknown keys with HTTP 400. Only `cost`, `latency`, and `vendorLockIn` are accepted. Example: `{ "cost": 0.4, "latency": 0.4, "vendorLockIn": 0.2 }`. Adding any extra key (e.g. `"reliability": 0.1`) returns 400. Values need not sum to 1.0 — the engine normalises them automatically. **Billing caveat — read before comparing providers**: All `costPerHour` and `totalCostPerHour` figures in explore results are simulated list-rate estimates. Modeled does not mean billing-validated — committed-use discounts, support surcharges, and marketplace fees are not included. Real invoices can differ materially. Treat these figures as a comparative shortlist, not a billing forecast. **`estimateDisclaimer` field**: Completed exploration results include an `estimateDisclaimer` string on `GET /multi-cloud/jobs/{jobId}/results`, `GET /multi-cloud/jobs/{jobId}` (when status is `completed`), and the UI preview endpoints. This is a static fidelity reminder: explore figures are flight-simulator-style planning estimates, not actual billing data. They reflect baseline on-demand rates under typical load. Internet egress is estimated and included in `totalCostPerHour` via the `egressCostPerHour` field in each strategy's `costBreakdown`. Still excluded: free-tier and Always-Free egress allowances, bandwidth pack credits (e.g. DigitalOcean's included TB/month), monitoring and support surcharges, sustained-peak ACU ramp-up for serverless tiers, free-tier and Always-Free credit offsets, and reserved-instance / committed-use / spot discounts. Allocations with different `skuClass` labels are not directly cost-comparable; see `crossClassNote` on any strategy that mixes non-equivalent product tiers. Actual spend can differ materially. Validate top candidates by materialising them as owned simulations and stepping at production RPS before making infrastructure commitments. **SLO constraint fields** (both optional, rank-last rather than filter): - `maxCostPerHour` (number, > 0): Budget ceiling in USD/hr. Strategies whose `totalCostPerHour` exceeds this value are moved to the tail of the ranked list (not filtered out), so you can see both in-budget and over-budget options in a single response. - `errorBudgetPct` (number, 0–100): SLO error-budget percentage. The engine models an implied error rate per provider count (single-provider ~0.50%, two-provider ~0.15%, three+ ~0.05%). Strategies whose implied rate exceeds this threshold are ranked last. Setting `0.1` effectively demotes all single-provider strategies below multi-provider alternatives. **SSE event types** (`GET /multi-cloud/jobs/{jobId}/stream`): the stream emits named events you can subscribe to individually: | Event | Payload | When | |---|---|---| | `init` | `{ jobId, status, progress, strategiesGenerated }` | Once, immediately on connect (reflects current job state). | | `strategy` | a single raw strategy object | For each strategy already accumulated, then for each new one as it is generated. | | `strategies-generated` | `{ strategiesGenerated, progress }` | Whenever the running strategy count changes — a dedicated live counter signal, emitted independently of `progress` so it stays current during the 0→40% generation phase even when the progress percentage hasn't ticked. | | `progress` | `{ progress, strategiesGenerated }` | Whenever the progress percentage changes. | | `complete` | `{ status, progress, strategiesGenerated, comparisonReport, latencyWarning, completedAt, error }` | Once, when the job reaches a terminal state (`completed`, `failed`, or `cancelled`); the stream then closes. | **Result payload includes**: Per-variant cost breakdown, latency estimates, vendor lock-in score, availability model, migration complexity rating, and multi-objective optimization ranking. DigitalOcean is a first-class candidate and typically produces the lowest-cost strategies. **Explore estimates vs. live /step values — read this before committing to a strategy**: The explore engine produces comparative estimates, not guaranteed runtime metrics. Two fields in particular require care: - `cost_usd_hr` in explore results — for serverless tiers (e.g. Aurora Serverless v2, Cloud Run, Lambda) this reflects the **ACU-floor / minimum-billing cost** at near-zero load, not the cost you will see at production RPS. At peak traffic the actual cost will be higher once ACUs ramp up. For provisioned tiers the figure is the baseline on-demand hourly rate without data-transfer, monitoring, or support surcharges. - `avgLatencyMs` in explore results — this is a **low-traffic baseline estimate** derived from provider region benchmarks, not a P95 or P99 at your production request rate. Under high concurrency, actual latency will be higher due to connection-pool pressure, cold-start ramp, and network jitter. **Recommendation**: treat explore output as a shortlist, not a final answer. Once you identify the top one or two strategies, materialise them as owned simulations (`POST /simulations`), apply your production traffic pattern, and step at target RPS to observe `metrics.latencyP95`, `metrics.costPerHour`, and `metrics.connectionPressure` under realistic load before committing. **`vendorLockInScore` direction**: In all explore results and comparison reports, `vendorLockInScore` is a **lock-in score** 0–100 where **higher = more locked in** (harder to migrate away). Single-provider strategies score 55–84 depending on how proprietary the DB service is (Aurora Serverless scores ~80, standard RDS ~62); multi-provider strategies score 20–50 depending on traffic distribution and provider count. A score of 75 for a single-provider AWS Aurora Serverless strategy means high lock-in; a score of 30 for a balanced AWS/GCP mix means low lock-in. Lower scores are preferable if portability is a goal. **New response fields** (available in `strategy.metrics` and per-strategy objects): - `p95LatencyMs` (number): Estimated P95 latency in milliseconds, computed as `avgLatencyMs × 1.5`. Use this to check whether a strategy's tail latency fits within a strict SLA before materialising it as a simulation. - `suggestedResources` (array of objects): Materializable SKU objects derived from each allocation. Each entry has `{provider, resourceType, size, count, hourlyRate}` — e.g. `{provider:"aws", resourceType:"compute", size:"t3.medium", count:4, hourlyRate:0.096}`. Use these to pre-populate infrastructure manifests or Terraform modules directly from the explore result. - `securityRecommendations` (array of strings): WAF/firewall palette IDs recommended for each provider footprint (e.g. `"aws-waf"`, `"gcp-cloud-armor"`, `"azure-waf"`, `"oci-waf"`). Separate from `suggestedResources` to keep SKU objects and security identifiers cleanly distinct. **SKU equivalence labels and cross-class notes** (on `strategy.allocations[]` and `strategy`): Each allocation carries three classification fields. `skuClass` is the primary label (compute wins when non-default; see precedence below). `computeSkuClass` and `databaseSkuClass` expose each dimension independently so cross-class mismatches are never masked. Allocations that share the same `skuClass` are cost-comparable at the class level; those with different values are not directly comparable. **Precedence for `skuClass`**: if `computeSkuClass` is `bare-metal-compute`, `free-tier-compute`, or `do-droplet`, that value is used as `skuClass` even when a `databaseSkuClass` is also present. For standard dedicated-vCPU compute (`general-compute`), the database class is promoted when set; the final fallback is `general-compute`. The eight class values are: | class | Dimension | Examples | |---|---|---| | `general-compute` | compute | Standard dedicated-vCPU VMs — EC2, Compute Engine, Azure VMs, OCI VM.Standard/Flex | | `do-droplet` | compute | DigitalOcean Droplets — shared-vCPU cloud VMs; different resource model from dedicated-vCPU instances | | `free-tier-compute` | compute | Always-free / zero-cost compute — OCI Always Free AMD micro VMs (costs excluded from totals) | | `bare-metal-compute` | compute | Dedicated bare-metal servers — OCI BM.Standard3.64, BM.Optimized3.36 | | `managed-rdbms` | database | Provisioned relational DBs — RDS, Cloud SQL, Azure SQL, … | | `oci-heatwave` | database | OCI MySQL HeatWave DB System and HeatWave Analytics Cluster — in-memory analytics accelerator; distinct billing model from managed-rdbms | | `serverless-db` | database | Pay-per-use / auto-scaling DBs — Aurora Serverless, Cloud Spanner, Autonomous DB, … | | `managed-nosql` | database | Wide-column / key-value NoSQL — Bigtable SSD/HDD, … | `crossClassNote` is present on a strategy when its allocations span more than one class in either the **compute** dimension (comparing `computeSkuClass` values) or the **database** dimension (comparing `databaseSkuClass` values). Each dimension is checked independently — a compute-tier mismatch is never masked by a shared database tier, and vice-versa. Absent when all allocations share the same class in both dimensions. **Simulation fidelity framing**: All cost and latency figures produced by the explore engine are flight-simulator-style planning estimates, not actual billing data. They are calibrated against public pricing pages (see `/fidelity`). Internet egress is estimated and included in `totalCostPerHour` via `egressCostPerHour`. Still excluded: free-tier and Always-Free egress allowances, bandwidth pack credits (e.g. DigitalOcean's included TB/month), monitoring/support surcharges, free-tier and Always-Free credit offsets, reserved-instance / committed-use / spot discounts, and the ACU ramp-up cost that serverless tiers incur at peak load. The `skuClass` label lets you quickly spot when two strategies are comparing non-equivalent tiers (e.g. a bare-metal OCI strategy vs a standard Droplet strategy, or provisioned RDS vs Aurora Serverless) before committing to a cost-difference figure. The `estimateDisclaimer` field in the explore result reflects these caveats for inline reference. **Aurora Serverless v2 as a cost-aware database candidate**: When exploring AWS-anchored strategies, the engine considers `aurora-serverless` (Aurora Serverless v2) as a low-cost database substitute for provisioned `aurora-postgresql` instances. Aurora Serverless v2 scales ACUs (Aurora Capacity Units) continuously between `minCapacity` and `maxCapacity`, so cost tracks actual load rather than peak provisioned size. It is the recommended choice when the workload is variable or has long idle periods. Exploration variants that include `aurora-serverless` will show a lower `cost_usd_hr` estimate compared with equivalent fixed-instance Aurora variants, along with a higher `scalability` rating and a lower `vendor_lock_in` penalty than proprietary NoSQL alternatives. **Cold-start trade-off**: ACU ramp-up from near-zero capacity can add 1–5 s of latency on the first request after an idle period; do not select `aurora-serverless` for latency-sensitive workloads (e.g. real-time APIs with strict P99 SLAs) unless a `minCapacity` ≥ 1 ACU is configured to keep the cluster warm. **Example: AWS architecture variant substituting provisioned Aurora with Aurora Serverless v2** Input simulation uses a provisioned `aurora-postgresql` instance. The exploration engine emits a cost-optimized AWS variant in the results: ```json { "variantId": "aws-cost-optimized", "provider": "aws", "resources": [ { "type": "compute", "provider": "aws", "name": "api-server", "size": "t3.medium" }, { "type": "database", "provider": "aws", "name": "primary-db", "serviceFamily": "aurora-serverless", "minCapacity": 0.5, "maxCapacity": 8 } ], "estimatedCostUsdHr": 0.09, "notes": "Aurora Serverless v2 replaces the provisioned aurora-postgresql instance; ACUs scale to near-zero during idle windows, reducing cost by ~40% for variable workloads." } ``` **Important**: `serviceFamily` shown in explore *result* objects (like `"serviceFamily": "aurora-serverless"` above) is part of the variant descriptor returned by the engine — it is **not** a field inside the `workloadProfile` request body. The `WorkloadProfile` schema does not include `serviceFamily`; any such property placed inside `workloadProfile` is silently ignored. The explore engine automatically considers Aurora Serverless v2 alongside provisioned Aurora instances for AWS-anchored strategies — submit a `workloadProfile` describing your compute instances, traffic, and latency requirements, and the engine generates both serverless and provisioned variants without any extra input field. **Side-by-side cost comparison — provisioned `aurora-postgresql` vs `aurora-serverless`**: The exploration engine returns both variants in the same results payload so agents can compare directly. Example response excerpt: ```json [ { "variantId": "aws-provisioned-db", "provider": "aws", "resources": [ { "type": "compute", "provider": "aws", "name": "api-server", "size": "t3.medium" }, { "type": "database", "provider": "aws", "name": "primary-db", "serviceFamily": "aurora-postgresql", "instanceClass": "db.r6g.large" } ], "estimatedCostUsdHr": 0.15, "notes": "Fixed-capacity provisioned Aurora PostgreSQL; cost is constant regardless of actual load." }, { "variantId": "aws-serverless-db", "provider": "aws", "resources": [ { "type": "compute", "provider": "aws", "name": "api-server", "size": "t3.medium" }, { "type": "database", "provider": "aws", "name": "primary-db", "serviceFamily": "aurora-serverless", "minCapacity": 0.5, "maxCapacity": 8 } ], "estimatedCostUsdHr": 0.09, "notes": "Aurora Serverless v2 scales ACUs to near-zero during idle windows. ~40% cost reduction versus the provisioned variant for variable workloads." } ] ``` Key fields for cost-aware selection: `estimatedCostUsdHr` is the per-hour estimate at average (not peak) load. For workloads with long idle periods or high variance, `aurora-serverless` will show a materially lower figure than `aurora-postgresql` with a fixed instance class. Prefer `aurora-serverless` when `estimatedCostUsdHr` is at least 20% lower and the workload is not latency-sensitive during cold ACU ramp-up. **GCP database tier — Cloud SQL, Cloud Spanner, Bigtable SSD, and Bigtable HDD**: When exploring GCP-anchored strategies, the engine considers four database service families as candidates: `cloud-sql` (provisioned PostgreSQL/MySQL instance), `cloud-spanner` (fully managed, horizontally scalable, globally distributed), `bigtable-ssd` (wide-column NoSQL, SSD-backed, low-latency analytics), and `bigtable-hdd` (wide-column NoSQL, HDD-backed, cost-optimised batch). Cloud SQL is lowest cost for single-region, schema-heavy, transactional workloads. Cloud Spanner scales horizontally without sharding and offers 99.999% multi-region SLA at higher base cost. Bigtable SSD delivers single-digit millisecond reads at scale for time-series and IoT workloads. Bigtable HDD is ~74% cheaper per node than SSD and is suited for batch analytics and archival access where latency tolerance exceeds ~50 ms. The engine surfaces all four variants so agents can evaluate the cost/capability trade-off directly. **Side-by-side cost comparison — `cloud-sql` vs `cloud-spanner` vs `bigtable-ssd` vs `bigtable-hdd`**: The exploration engine returns all variants in the same results payload so agents can compare directly. Example response excerpt: ```json [ { "variantId": "gcp-cloud-sql", "provider": "gcp", "resources": [ { "type": "compute", "provider": "gcp", "name": "api-server", "size": "n2-standard-2" }, { "type": "database", "provider": "gcp", "name": "primary-db", "serviceFamily": "cloud-sql", "instanceClass": "db-n1-standard-2" } ], "estimatedCostUsdHr": 0.12, "notes": "Provisioned Cloud SQL PostgreSQL instance (2 vCPU, 7.5 GB). Fixed cost regardless of load; well-suited for small, steady-state, single-region workloads." }, { "variantId": "gcp-cloud-spanner", "provider": "gcp", "resources": [ { "type": "compute", "provider": "gcp", "name": "api-server", "size": "n2-standard-2" }, { "type": "database", "provider": "gcp", "name": "primary-db", "serviceFamily": "cloud-spanner", "processingUnits": 100 } ], "estimatedCostUsdHr": 0.28, "notes": "Cloud Spanner at 100 processing units (0.1 node). ~2.3x the cost of Cloud SQL at small scale, but scales horizontally without sharding, offers 99.999% multi-region SLA, and eliminates manual capacity planning as write throughput grows." }, { "variantId": "gcp-bigtable-ssd", "provider": "gcp", "resources": [ { "type": "compute", "provider": "gcp", "name": "api-server", "size": "n2-standard-2" }, { "type": "database", "provider": "gcp", "name": "primary-db", "serviceFamily": "bigtable-ssd", "nodes": 1 } ], "estimatedCostUsdHr": 0.74, "notes": "Cloud Bigtable SSD node ($0.65/node-hr). Wide-column NoSQL with single-digit ms latency at millions of rows/sec. Best for time-series, IoT telemetry, and low-latency analytics; no SQL support." }, { "variantId": "gcp-bigtable-hdd", "provider": "gcp", "resources": [ { "type": "compute", "provider": "gcp", "name": "api-server", "size": "n2-standard-2" }, { "type": "database", "provider": "gcp", "name": "primary-db", "serviceFamily": "bigtable-hdd", "nodes": 1 } ], "estimatedCostUsdHr": 0.26, "notes": "Cloud Bigtable HDD node ($0.17/node-hr). ~74% cheaper than SSD; suited for batch analytics and archival workloads with p99 latency tolerance above ~50 ms." } ] ``` Key fields for cost-aware selection: at small, steady-state load `cloud-sql` will show a materially lower `estimatedCostUsdHr` than `cloud-spanner`. Prefer `cloud-spanner` when the workload requires horizontal write scaling, multi-region strong consistency, or the projected growth path would require manual sharding of a Cloud SQL instance. Prefer `cloud-sql` when `estimatedCostUsdHr` savings exceed 50% and the workload fits within a single provisioned instance with no cross-region consistency requirement. Prefer `bigtable-ssd` for high-throughput NoSQL workloads (IoT, time-series, recommendations) where p99 read latency must stay below 10 ms. Prefer `bigtable-hdd` when the same NoSQL data model is acceptable but the workload is batch-oriented or archival and a ~74% cost reduction per node outweighs the higher latency. **Azure database tier — Azure SQL vs Cosmos DB**: When exploring Azure-anchored strategies, the engine considers both `azure-sql` (provisioned relational database, SQL Server-compatible) and `cosmos-db` (globally distributed, multi-model, serverless or provisioned throughput) as database candidates. Azure SQL is lower cost for single-region, schema-heavy, transactional workloads with predictable concurrency at moderate throughput. Cosmos DB charges per Request Unit per second (RU/s): at its 400 RU/s minimum the entry cost is low, but RU/s provisioning scales linearly — at high write concurrency Cosmos DB can exceed Azure SQL cost significantly. Cosmos DB offers single-digit-millisecond latency at any scale, turnkey global distribution across Azure regions, and multiple consistency levels — making it the preferred choice when the workload requires low-latency reads at global scale, schema flexibility, or multi-region active-active writes. The engine surfaces both variants so agents can evaluate the cost/capability trade-off directly. **Side-by-side cost comparison — `azure-sql` vs `cosmos-db`**: The exploration engine returns both variants in the same results payload so agents can compare directly. Example response excerpt: ```json [ { "variantId": "azure-azure-sql", "provider": "azure", "resources": [ { "type": "compute", "provider": "azure", "name": "api-server", "size": "Standard_D2s_v3" }, { "type": "database", "provider": "azure", "name": "primary-db", "serviceFamily": "azure-sql", "tier": "GeneralPurpose", "vCores": 2 } ], "estimatedCostUsdHr": 0.18, "notes": "Provisioned Azure SQL General Purpose, 2 vCores. Fixed cost regardless of request volume; well-suited for relational, single-region workloads with predictable concurrency and structured schemas." }, { "variantId": "azure-cosmos-db", "provider": "azure", "resources": [ { "type": "compute", "provider": "azure", "name": "api-server", "size": "Standard_D2s_v3" }, { "type": "database", "provider": "azure", "name": "primary-db", "serviceFamily": "cosmos-db", "throughputModel": "provisioned", "requestUnitsPerSecond": 400 } ], "estimatedCostUsdHr": 0.048, "notes": "Cosmos DB provisioned at 400 RU/s (minimum). Lower entry cost at minimal throughput, but RU/s charges scale linearly with required throughput — at high write volumes Cosmos DB can exceed Azure SQL cost. Offers multi-model API (SQL, MongoDB, Cassandra), tunable consistency, and turnkey multi-region replication not available in Azure SQL." } ] ``` Key fields for cost-aware selection: at low throughput `cosmos-db` entry cost (400 RU/s minimum) can appear lower than `azure-sql`, but RU/s charges scale linearly — at sustained high write concurrency `azure-sql` will show a materially lower `estimatedCostUsdHr`. Prefer `cosmos-db` when the workload requires global distribution, sub-10 ms reads at scale, schema flexibility, or multi-region active-active writes. Prefer `azure-sql` when the workload is single-region, relational, and concurrency is predictable — the fixed vCore pricing model becomes significantly cheaper once Cosmos DB RU/s provisioning would need to exceed ~2000 RU/s to serve the load. **Azure SQL Managed Instance tiers — MI_GP_Gen5_4 vs MI_BC_Gen5_4**: SQL Managed Instance (SQL MI) is a fully managed deployment option with near 100% SQL Server compatibility, targeting lift-and-shift migrations. The simulator tracks two committed reference shapes for East US (June 2026): | Shape | SKU | vCores | RAM | Committed rate | Monthly approx. | |---|---|---|---|---|---| | MI_GP_Gen5_4 | General Purpose, Gen5, 4 vCore | 4 | 20.4 GB | $0.741/hr | ~$541/mo | | MI_BC_Gen5_4 | Business Critical, Gen5, 4 vCore | 4 | 20.4 GB | $1.906/hr | ~$1,391/mo | The General Purpose tier uses remote storage (Azure Premium storage, ~5.5 ms I/O latency) and is suited for workloads that can tolerate occasional storage I/O latency. The Business Critical tier provisions local SSD storage with a built-in read-scale secondary replica, delivering sub-millisecond I/O and built-in high availability at roughly 2.6× the General Purpose cost. Choose General Purpose for cost-sensitive, write-heavy OLTP workloads where remote storage latency is acceptable. Choose Business Critical when the workload has strict P99 latency requirements (e.g. ≤ 5 ms reads), needs an in-region read replica for reporting without extra licensing cost, or requires maximum availability (BC provides a built-in local secondary replica that survives storage failures, giving higher effective resilience than GP's remote-storage model despite both tiers publishing a 99.99% uptime SLA). Both shapes are tracked in the weekly pricing drift check (`azure-sql-mi-gp-gen5-4` and `azure-sql-mi-bc-gen5-4`) and appear in the Fidelity benchmark status panel alongside Azure SQL Database and Cosmos DB entries. **OCI database tier — Autonomous Transaction Processing (ATP) vs MySQL HeatWave (with and without HeatWave cluster)**: When exploring OCI-anchored strategies, the engine considers three database variants: `autonomous-db` (Oracle Autonomous Transaction Processing, self-driving, self-securing, ECPU-billed), `mysql-heatwave` (managed MySQL DB System on E4.Flex compute, no cluster — pure OLTP), and `mysql-heatwave-cluster` (same DB System plus one in-memory HeatWave cluster node for HTAP analytics). ATP is the higher-cost option at entry level but bundles Oracle-managed patching, in-memory columnar query acceleration, automatic index tuning, and serverless ECPU scaling — making it well-suited for mixed OLTP/analytics workloads that need zero-DBA management and predictable SLAs. MySQL HeatWave (DB System without the HeatWave cluster) is significantly cheaper for pure transactional MySQL workloads: at 2 OCPU / 16 GB on E4.Flex the all-in hourly rate is materially lower than ATP's 2-ECPU baseline, and the familiar MySQL wire protocol avoids any application re-platforming cost. Adding a HeatWave cluster node unlocks in-database analytics on live OLTP data without ETL at an additional $0.0255/node-hr. The engine surfaces all three variants so agents can evaluate the full cost/capability spectrum directly. **Side-by-side cost comparison — `autonomous-db` vs `mysql-heatwave` vs `mysql-heatwave-cluster`**: The exploration engine returns all three variants in the same results payload so agents can compare directly. Example response excerpt: ```json [ { "variantId": "oci-autonomous-db", "provider": "oci", "resources": [ { "type": "compute", "provider": "oci", "name": "api-server", "size": "VM.Standard.E4.Flex" }, { "type": "database", "provider": "oci", "name": "primary-db", "serviceFamily": "autonomous-db", "tier": "standard", "ecpus": 2 } ], "estimatedCostUsdHr": 0.789, "notes": "OCI Autonomous AI Transaction Processing (Standard), 2 ECPUs at $0.336/ECPU-hr plus blended storage (~$0.117/hr). OCI rebranded Autonomous Database as Autonomous AI (part B95702). Self-driving: automated patching, tuning, and scaling. Built-in in-memory columnar processing for mixed OLTP/analytics. Best when zero-DBA operation, Oracle-compatible SQL, or HTAP capability is required." }, { "variantId": "oci-mysql-heatwave", "provider": "oci", "resources": [ { "type": "compute", "provider": "oci", "name": "api-server", "size": "VM.Standard.E4.Flex" }, { "type": "database", "provider": "oci", "name": "primary-db", "serviceFamily": "mysql-heatwave", "shape": "MySQL.VM.Standard.E4.Flex", "ocpus": 2, "memoryGb": 16 } ], "estimatedCostUsdHr": 0.074, "notes": "OCI MySQL HeatWave DB System, 2 OCPU / 16 GB on E4.Flex — no HeatWave cluster. $0.025/OCPU-hr × 2 + $0.0015/GB-hr × 16 = $0.074/hr. Standard MySQL wire protocol — zero application re-platforming. Lowest-cost OCI database option for pure transactional MySQL workloads. To add in-database HTAP analytics on live OLTP data, see the `oci-mysql-heatwave-cluster` variant below ($0.0995/hr with 1 cluster node)." }, { "variantId": "oci-mysql-heatwave-cluster", "provider": "oci", "resources": [ { "type": "compute", "provider": "oci", "name": "api-server", "size": "VM.Standard.E4.Flex" }, { "type": "database", "provider": "oci", "name": "primary-db", "serviceFamily": "mysql-heatwave", "shape": "MySQL.VM.Standard.E4.Flex", "ocpus": 2, "memoryGb": 16 }, { "type": "database-addon", "provider": "oci", "name": "heatwave-cluster", "serviceFamily": "mysql-heatwave-cluster", "nodes": 1, "shape": "HeatWave.512GB" } ], "estimatedCostUsdHr": 0.0995, "notes": "OCI MySQL HeatWave DB System + 1 HeatWave cluster node for HTAP (hybrid transactional/analytical processing). DB System: $0.074/hr (2 OCPU / 16 GB, same as `oci-mysql-heatwave`). HeatWave cluster node: $0.0255/node-hr. Combined: $0.074 + $0.0255 = $0.0995/hr. Enables in-database analytics on live OLTP data without ETL or a separate analytics tier. At $0.0995/hr the combined cost is comparable to `autonomous-db` ($0.789/hr) — choose based on MySQL wire-protocol compatibility requirements and in-database analytics throughput needs." } ] ``` Key fields for cost-aware selection: `autonomous-db` carries roughly twice the hourly cost of a baseline `mysql-heatwave` DB System at equivalent scale, but that premium covers automated DBA operations, Oracle SQL compatibility, and built-in HTAP acceleration. Adding a HeatWave cluster node to the MySQL DB System (`mysql-heatwave-cluster`) pushes the all-in cost to approximately $0.0995/hr — the cluster node add-on ($0.0255/node-hr) is modest and the combined cost is comparable to `autonomous-db` ($0.789/hr). Prefer `autonomous-db` when the workload requires Oracle-compatible SQL, mixed OLTP/analytics queries without a separate analytics tier, or zero-ops database management with automatic index and performance tuning. Prefer `mysql-heatwave` (no cluster) when the application is MySQL-native, the workload is purely transactional, and the cost saving versus ATP exceeds the value of Oracle's self-managing features — the MySQL wire protocol means no driver or ORM changes are needed on migration. Choose `mysql-heatwave-cluster` when MySQL wire-protocol compatibility is a hard requirement **and** in-database analytics on live OLTP data is the primary workload driver; at $0.0995/hr it is a cost-competitive alternative to `autonomous-db` ($0.789/hr) for HTAP workloads. **OCI Block Volume — Ultra High Performance (UHP) tiers**: OCI Block Volume UHP scales linearly from 30 to 120 VPUs/GB. All four tiers share the same volume-level caps (225,000 IOPS, 2,680 MB/s) — the VPU count controls per-GB density, not the absolute caps. Higher VPU tiers saturate those caps on smaller volumes, making them cost-efficient for high-IOPS-density workloads where volume size is constrained. Pricing is based on a ~430 GB reference volume (us-ashburn-1, May 2026). ```json [ { "size": "ultra-high-performance-30vpus", "provider": "oci", "serviceFamily": "block-volume", "estimatedCostUsdHr": 0.041, "iopsMax": 225000, "throughputMbsMax": 2680, "iopsPerGb": 90, "throughputKbpsPerGb": 900, "notes": "OCI Block Volume Ultra High Performance, 30 VPUs/GB. $0.0255/GB-mo base + 20 extra VPUs × $0.00225/VPU/GB-mo = $0.0705/GB-mo; flat sim rate $0.041/hr (~430 GB reference). 90 IOPS/GB and 900 KBPS/GB; volume caps: 225,000 IOPS and 2,680 MB/s. IOPS cap saturates at ~2,500 GB. Baseline UHP tier — lowest per-GB VPU cost; suitable for large volumes with moderate IOPS density." }, { "size": "ultra-high-performance-60vpus", "provider": "oci", "serviceFamily": "block-volume", "estimatedCostUsdHr": 0.082, "iopsMax": 225000, "throughputMbsMax": 2680, "iopsPerGb": 180, "throughputKbpsPerGb": 1800, "notes": "OCI Block Volume Ultra High Performance, 60 VPUs/GB. $0.0255/GB-mo base + 50 extra VPUs × $0.00225/VPU/GB-mo = $0.138/GB-mo; flat sim rate $0.082/hr (~430 GB reference). 180 IOPS/GB and 1,800 KBPS/GB; volume caps unchanged at 225,000 IOPS and 2,680 MB/s. IOPS cap saturates at ~1,250 GB — half the volume size needed vs. 30 VPUs to reach peak IOPS." }, { "size": "ultra-high-performance-90vpus", "provider": "oci", "serviceFamily": "block-volume", "estimatedCostUsdHr": 0.123, "iopsMax": 225000, "throughputMbsMax": 2680, "iopsPerGb": 270, "throughputKbpsPerGb": 2700, "notes": "OCI Block Volume Ultra High Performance, 90 VPUs/GB. $0.0255/GB-mo base + 80 extra VPUs × $0.00225/VPU/GB-mo = $0.2055/GB-mo; flat sim rate $0.123/hr (~430 GB reference). 270 IOPS/GB and 2,700 KBPS/GB; volume caps unchanged at 225,000 IOPS and 2,680 MB/s. IOPS cap saturates at ~834 GB. Throughput per-GB rate (2,700 KBPS/GB) nearly matches the 2,680 MB/s cap, so throughput cap is reached at ~993 GB." }, { "size": "ultra-high-performance-120vpus", "provider": "oci", "serviceFamily": "block-volume", "estimatedCostUsdHr": 0.163, "iopsMax": 225000, "throughputMbsMax": 2680, "iopsPerGb": 360, "throughputKbpsPerGb": 3600, "notes": "OCI Block Volume Ultra High Performance, 120 VPUs/GB. $0.0255/GB-mo base + 110 extra VPUs × $0.00225/VPU/GB-mo = $0.273/GB-mo; flat sim rate $0.163/hr (~430 GB reference). 360 IOPS/GB and 3,600 KBPS/GB; volume caps unchanged at 225,000 IOPS and 2,680 MB/s. IOPS cap saturates at ~625 GB — the most IOPS-dense UHP tier. Both caps are reachable on relatively small volumes, maximising peak performance per GB." } ] ``` Key fields for tier selection: all UHP tiers share the same absolute volume caps (225,000 IOPS, 2,680 MB/s). VPU count determines the per-GB density at which those caps are reached. Choose a higher VPU tier when volumes are small but IOPS demand is high — a 625 GB volume at 120 VPUs delivers the same peak IOPS as a 2,500 GB volume at 30 VPUs, at a higher per-GB price but with less provisioned capacity. Prefer 30 VPUs for large volumes where per-GB rate is the primary cost driver and the IOPS cap is reachable purely through volume size. --- ### x402 Pay-Per-Request API Anonymous pay-per-request access using USDC-on-Base. No API key or account required — attach an `X-PAYMENT` header and call any of the 41 metered endpoints. Use `GET /billing/x402/config` first to discover whether x402 is enabled and to read exact USDC amounts. The paid call types include: `simulation_create`, `simulation_list`, `simulation_get`, `simulation_cost_breakdown`, `simulation_step_hybrid`, `rl_env_create`, `rl_env_get`, `rl_env_delete`, `rl_env_reset`, `rl_env_observation`, `rl.step`, `rl.batch_step`, `rl.eval`, `ai.explain`, `ai.optimize`, `ai.troubleshoot`, `ai_bottleneck`, `ai_analysis`, `simulation.inject_traffic`, `simulation.inject_failure`, `simulation.inject_failure_create`, `simulation.inject_failure_update`, `simulation.inject_failure_delete`, `simulation.resize`, `simulation.apply_right_sizing`, `right_sizing_hint`, `validate_cost_accuracy`, `validate_performance_accuracy`, `benchmark.validate`, `wallet_session`, `rl_eval_status`, `multicloud.status`, `multicloud_results`, `multicloud_partial_results`, `multicloud_stream`, `ai.status` ($0.0010 each), and `chaos.run`, `chaos_batch`, `multicloud.explore`, `optimization.run`, `prediction.validate`, `prediction.optimize_thresholds` ($0.0050 each). | Method | Path | Auth | Description | |---|---|---|---| | GET | /billing/x402/config | Public | x402 configuration: enabled flag, payTo wallet address, network, USDC asset address, facilitator URL, credits-per-USDC rate, and per-call-type price table | | GET | /billing/x402/balance | Public | Credit balance for a wallet address (`?address=0x...`); returns 404 if the address has never paid for a call | | GET | /billing/x402/transactions | Public | Payment transaction history for a wallet address (`?address=0x...&limit=50`); positive delta = payment received, negative delta = call deducted | | POST | /billing/x402/settlements/{reconciliationId}/reconcile | Public | Reconcile a previously pending facilitator settlement using its opaque reconciliation handle; no API key or x402 payment is required. Returns confirmed settlement (200), still-pending settlement (202), invalid handle/not found (404), invalid or rejected settlement (402), or facilitator-query failure (502) | --- ## Webhook / Callback System Asynchronous notifications POSTed to agent-provided URLs on job completion. - **Supported jobs**: `analysis/optimize`, `chaos/run`, `chaos/batch`, `predictions/validate`, `predictions/optimize-thresholds`, `multi-cloud/explore`, `rl/environments` (episode completion). - **Security**: HTTPS-only destinations; SSRF protection (private IP ranges and localhost are blocked). - **Signature**: Every webhook includes an `X-Webhook-Signature` header containing an HMAC-SHA256 signature. Verify by computing `HMAC-SHA256(webhookSecret, rawBody)` and comparing with constant-time equality. - **Retry logic**: Up to 3 delivery attempts with exponential backoff (0 s, 2 s, 8 s). Each attempt has a 10-second timeout. - **Delivery tracking**: Job responses include `webhookDeliveryStatus` (`"pending"`, `"delivered"`, or `"failed"`), `webhookDeliveryAttempts`, `webhookDeliveryError`, and `webhookDeliveredAt`. - **Payload format**: JSON with `{ event, jobId, status, data, timestamp }`. The `event` field follows the pattern `".completed"` or `".failed"`. --- ## Authentication & Security Notes - All authenticated endpoints require `Authorization: Bearer ` header. - API keys are bcrypt-hashed at rest; the raw key is only shown once at issuance. - Keys support scope-based access control: `read`, `write`, `admin`. - Rate limiting is enforced per key (default: 1000 requests/hour). - Exhausting credits blocks further simulation steps (`credits_exhausted` event is fired). - Webhook URLs undergo SSRF validation; private IP ranges and DNS rebinding are blocked. --- ## Integration & Tech Stack - **AI**: OpenAI GPT-5 via Replit AI Integrations - simulation explanations, bottleneck analysis, architecture optimization suggestions, and beginner-friendly explanations. - **Payments**: Stripe - credit top-ups via Stripe Checkout; webhook handler for authoritative payment confirmation (`topup_completed` event). - **Analytics**: PostHog - product analytics, session recordings, and feature flags (no-op when key is absent). - **Database**: PostgreSQL via Drizzle ORM and Neon serverless adapter (in-memory in development). - **Frontend**: React 18, TypeScript, Vite, Wouter, TanStack Query, Shadcn/ui, Tailwind CSS. - **Backend**: Express.js, TypeScript, Node.js. --- ## MCP Demo Tools (Smithery) Nine anonymous tools are served via the Smithery-hosted MCP server (`smithery.ai/server/@canvas-cloud-ai/cwm`). No API key or account is required — connect and call immediately. The hosted `/mcp` endpoint exposes all 62 tools when initialized with `x-api-key: ` or `Authorization: Bearer `; send the same key on subsequent session requests. Smithery can also forward a bare `cwm_…` key in `Authorization`. Do not send both headers. Invalid, malformed, or unsupported key-like headers return HTTP 401, never the anonymous demo. The paid stateless one-call operation is intentionally REST-only at `POST /api/simulations/stateless`. MCP tool invocation does not preserve the HTTP 402 challenge, selected Base/Solana payment header, settlement response, and exact-once replay boundary end to end. No MCP tool was added, and published MCP tool counts remain unchanged. ### Anonymous session limits | Cap | Value | |---|---| | Simulations per session | 3 | | Steps per simulation | 20 | | Max traffic (RPS) | 1 000 | ### Demo tool reference (9 anonymous tools) | Tool | Description | Annotations | |---|---|---| | `api.spec` | Return the OpenAPI spec URL and format. No input required. | read-only, idempotent | | `scenario.list` | List compact cards from the live scenario library. Anonymous cards include only scenarios with at most 10 resources; authenticated cards include the complete catalog. Each card includes `id`, `title`/`name`, description, difficulty, tags, category, duration, provider summary, and resource/connection counts; topology and experiment graphs are omitted. No input required. | read-only, idempotent | | `scenario.get` | Hydrate one selected scenario by `scenarioId`. Returns the full graph, including `resources` and `connections`, plus optional seed, resilience, traffic/failure presets, and incident metadata. Use it to inspect or customize a graph, or pass the id directly to `simulation.create`. | read-only, idempotent | | `simulation.create` | Create a simulation. Accepts `name` (string) plus exactly one graph source: `scenarioId` (a live id from `scenario.list`, expanded server-side) or explicit `resources` / optional `connections` arrays. Do not send scenarioId with resources or connections. Scenario traffic and failure presets are not applied automatically. Also accepts optional `traffic` (RPS, default 0) and a simulation-wide CPU HPA target via `autoscalingTargetCpu`, `scaleOutCpuThreshold`, `scaleOutCpuPercent`, or `autoscaleTargetCpuPercent` (all supplied names must agree; a misnamed near-miss field nested under a resource's `characteristics` is rejected with 400). For a per-resource CPU target that overrides the simulation-wide value for just one resource, set `characteristics.scaleOutCpuThreshold` / `characteristics.scaleInCpuThreshold` on that resource instead. Returns `id` — store this for subsequent calls. | write, non-idempotent | | `simulation.step` | Advance the simulation by one time step. Input: `simulationId`. Returns `currentTime`, `metrics` (latencyP50/P95/P99, cpuUsage, throughput, errorRate, costPerHour), and `events`. | write, non-idempotent | | `simulation.metrics` | Read the current simulation state and full metrics history. Input: `simulationId`. Returns `simulation` (id, name, currentTime, traffic, resources with health/status) and `metrics` array. | read-only, idempotent | | `simulation.provider_api_limits` | Read-only bounded provider control-plane simulation. Use for provider quotas, throttling, retries, backoff, queueing, worker/concurrency limits, and control-plane operations—not `simulation.step` or `/api/simulations/stateless`. One service per request. Supported AWS EC2/ELBv1/ELBv2, Aurora Serverless v1 Data API (`rds-data-api-aurora-serverless-v1`, `database`, independent `requests-per-second` / `concurrent-requests` models: `aws.rds-data-api-aurora-serverless-v1.requests-per-second`, `aws.rds-data-api-aurora-serverless-v1.concurrent-requests`), and Azure ARM tuples resolve as `catalog_default`; see REST row for scopes and model details. IAM, generic RDS control-plane, Aurora Serverless v2 Data API, provisioned Aurora Data API, Auto Scaling, and S3 control-plane tuples need an account-scoped override or return `unsupported_policy`; S3 object per-prefix figures are data-plane, not a control-plane quota. A concurrency sweep shares a 1,000,000-event and 100,000-unit aggregate budget across candidates, the whole request has a 2-second wall-clock deadline, and oversized sweeps are rejected before scheduling with non-retryable HTTP 413 `PROVIDER_LIMIT_WORK_BUDGET`. Two executions may run concurrently. A third request returns non-retryable 409 `PROVIDER_LIMIT_BUSY`; an identical active or interrupted request returns non-retryable 409 `PROVIDER_LIMIT_DUPLICATE` with a five-minute guard. Returns scheduler results plus catalog/override provenance and as-of metadata; does not create resources, advance `/step`, call a cloud API, discover live quotas, or change simulation state. | read-only, idempotent | | `simulation.inject_traffic` | Set the simulation's traffic level. Input: `simulationId`, `targetRps` (absolute RPS, recommended) or `deltaPercent` (relative change, e.g. 75 for +75%); pass `{"random": true}` to opt into an uncontrolled random spike. Empty body `{}` returns 400. Server-caps at 1 000 RPS in demo mode. Returns updated `traffic` and `status`. | write, non-idempotent | | `simulation.inject_failure` | Inject a failure into a selected demo resource. Use `simulation.recover_resource` for reversible `instance_down`/database overload failures. | write, non-idempotent | | `simulation.recover_resource` | Start recovery for a failed resource using `resourceId` or `resourceName`. Returns recovery progress; call `simulation.step` until the state is `healthy`. | write, non-idempotent | | `simulation.delete` | Permanently delete an owned demo simulation, revoke its capability, clear the session pointer, and free a simulation slot. Returns `{deleted: true, id}`. | destructive write, idempotent | ### Typical agent demo loop 1. Call `scenario.list` — pick a compact scenario card and save its `id`. 2. Call `simulation.create` with `name` and `scenarioId` for server-side expansion, then save the returned simulation `id`. Call `scenario.get` first only when you need to inspect or customize the graph; larger scenarios require authentication. 3. Call `simulation.inject_traffic` with a target RPS (e.g. 500). 4. Loop: call `simulation.step` and read `metrics.cpuUsage`, `metrics.errorRate`, `metrics.costPerHour`. 5. Optionally call `simulation.inject_failure`, then `simulation.recover_resource` and keep stepping until recovery is healthy. 6. Call `simulation.metrics` for the full resource-health breakdown, then `simulation.delete` to intentionally free the demo slot. --- ## WebMCP Browser Tools AI agents navigating the site in a WebMCP-enabled browser (Chrome 149+ with the `#enable-webmcp` flag or a valid Origin Trial token) can call platform actions directly via `navigator.modelContext`. The registration is a pure progressive enhancement — non-WebMCP browsers are completely unaffected. The `Permissions-Policy: webmcp=*` response header is set on all pages to opt the site in to the API. ### Registered Tools | Tool name | Description | |---|---| | `run_simulation_step` | Advance the active simulation by one tick; returns updated CPU, latency, throughput, error rate, and cost snapshot. | | `reset_simulation` | Pause the active simulation and reset traffic to zero. | | `simulation.inject_traffic_pattern` | Inject a named traffic pattern (`ramp`, `burst`, `step`, `wave`, `custom`; legacy aliases `spike`→`burst`, `sine`→`wave`, `gradual_increase`→`ramp`) into the running simulation. | | `simulation.inject_failure` | Fail a specific node in the running simulation. Supply `resourceId` or `resourceName` to target a named resource (deterministic, recommended). Pass `{"random": true}` to explicitly opt into random node selection (non-deterministic). Empty body `{}` returns 400. | | `simulation.inject_failure_create` | Create a named, scheduled failure injection (`POST /simulations/{id}/failures`): specify type (`instance_kill`, `instance_down`, `az_outage`, `database_overload`, `network_latency`, `spot_interruption`), start/end steps, and optional target resource or zone. For `database_overload`, set top-level `severity` (minor/moderate/severe, default moderate) or `intensity` (0–1 mapped to severity), never inside `parameters` or both; the response echoes effective severity. The load-sensitive CPU, ACU, error and latency effects are modeled assumptions, not AWS guarantees; throughput is successful RPS. `spot_interruption` is the bounded AWS EKS migration lifecycle; `instance_kill` is permanent; `instance_down` is reversible. Requires `write` scope or wallet JWT. | | `simulation.inject_failure_update` | Update an existing failure injection (`PATCH /failures/{id}`): deactivate it early (`isActive: false`) or adjust its end time. Requires `write` scope or wallet JWT. | | `simulation.inject_failure_delete` | Permanently delete a failure injection (`DELETE /failures/{id}`). Returns 204 No Content. Requires `write` scope or wallet JWT. | | `get_current_metrics` | Read the latest CPU utilization, latency percentiles (P50/P95/P99), error rate, throughput, and estimated hourly cost for the active simulation. | | `navigate_to_page` | Navigate the browser to a named platform page: `home`, `scenarios`, `billing`, `docs`, `api-reference`. | All simulation tools accept an optional `simulationId` input; when omitted they resolve the active simulation from the URL query string (`?simulation=`) or by fetching the first simulation from the API. Tools call the `/api/ui/*` REST endpoints (no API key required) and follow the same rate limits as the browser workspace. --- ## Behavior-Model Provenance & GPU-Signal Warnings **OpenShift pricing provenance:** `normalizedConfig.resources[*].openshiftPricingStatus` describes the platform/license line independently from worker and control-plane charges. ROSA published platform lines are official across their supported AWS standard regions. The canonical ARO evidence is Microsoft's Azure Red Hat OpenShift pricing table: the East US D4s v3 (4 vCPU) OpenShift line is `$124.830/month`, and `$124.830/month ÷ 730 = $0.171/hour per 4 vCPU`. It was verified on 2026-08-27 at https://azure.microsoft.com/en-us/pricing/details/openshift/. The retest-reported $0.2565/hour is a valid topology-derived platform total for three 2-vCPU workers: `$0.171 × (6 effective worker vCPU ÷ 4 vCPU per billing unit) = $0.2565/hour`. It is not a new published rate or independent pricing source. The canonical accuracy comparison uses two 2-vCPU workers (4 effective worker vCPU, one billing unit), so its derived platform total is the same `$0.171/hour` as the published unit rate. The ARO license amount is an official published line only when the resource has `location.regionKey` set to Azure East US (accepted spellings include `East US` and `eastus`); a missing or different Azure region is reported as `estimated` because the published source is regional. Cost-breakdown rows retain `pricingSource`, `effectiveDate`, `referenceRegion`, `billingUnit`, `billingUnitVcpu`, `publishedUnitRatePerHour`, `effectiveWorkerCount`, `effectiveWorkerVcpu`, `derivedPlatformFeePerHour`, and per-pool `workerTopology`. They always keep `platform`, `control-plane`, and `worker` charges separate. For ARO, normalize comparisons with `derivedPlatformFeePerHour = publishedUnitRatePerHour × effectiveWorkerVcpu ÷ billingUnitVcpu`; do not compare a multi-node topology total directly with the one-unit published rate. Every entry in `normalizedConfig.resources` (returned on `GET /api/simulations/:id`, creation responses, and step-hybrid responses) carries a `behaviorModel` field that identifies the exact latency simulation path applied to that resource: | `latencyPath` | When used | |---|---| | `gpu-inference/ttft-decode` | Kubernetes cluster with `characteristics.inferenceMode: true` — token-economics TTFT+decode model | | `generic-kubernetes/load-saturation` | Kubernetes cluster without `inferenceMode` (or `inferenceMode: false`) — generic saturation model | | `generic-compute/load-saturation` | Compute (VM) resources | | `database/connection-saturation` | Database resources — connection-pool pressure model | | `storage/iops-saturation` | OCI Block Volume only (`provider: "oci"`, `serviceFamily: "block-volume"` or `maxIops` set) — IOPS ceiling tracked; latency and error-rate penalties propagated to co-located databases when ceiling is exceeded | | `generic/flat-rate` | All other storage (S3, GCS, Azure Blob, OCI Object Storage), network, cache, queue, security — cost-only, no latency simulation path | **GPU misconfiguration warning (`inferenceWarnings`):** When a Kubernetes resource carries GPU/ inference characteristics (`accelerator`, `tokensPerRequest`, `inputTokensPerRequest`) but `inferenceMode` is absent or false, a non-fatal `inferenceWarnings` string array is added to that resource's normalized entry. The warning names the detected signals, states which model will run, and gives the exact field to set to change it. Example: ```json ["Resource has inference/GPU characteristics (accelerator: h100, tokensPerRequest: 512) but inferenceMode is not enabled. The generic-kubernetes/load-saturation latency model will be used. Set characteristics.inferenceMode: true to activate gpu-inference/ttft-decode."] ``` **Low-agreement rule-engine path diagnostic (`behaviorMismatch`):** When a hybrid step produces `agreement.level === "low"` AND at least one Kubernetes resource has GPU signals without `inferenceMode`, the `hybridDecision` object carries a `behaviorMismatch` diagnostic. This is a factual note about the rule engine's latency path — not a claim about the ML predictor's internals (the hybrid ML predictor derives its latency prediction from the rule engine's baseline output): ```json { "detected": true, "cause": "Rule engine used generic-kubernetes/load-saturation for resource(s) that carry GPU/inference characteristics (accelerator, tokensPerRequest, or inputTokensPerRequest) but have inferenceMode: false. If this simulation is intended to model GPU inference workloads, setting characteristics.inferenceMode: true would activate the gpu-inference/ttft-decode latency model in the rule engine, which may change the rule/ML agreement score.", "affectedResources": [""] } ``` **Agent guidance:** If your hybrid simulation produces unexpectedly low agreement and the resources are GPU/LLM inference nodes, check `normalizedConfig.resources[*].inferenceWarnings` first. If warnings are present, it means the rule engine is using the generic K8s latency model instead of the GPU inference model. Add `"inferenceMode": true` to the resource's `characteristics` to activate `gpu-inference/ttft-decode` in the rule engine, which may resolve the discrepancy. --- ## Kubernetes normalization receipt and modeling limits Every create, get, and step response for a simulation includes an immutable `normalizationReceipt` (`gke-normalization-receipt/v2`). Use this receipt—not a natural-language prompt—to compare GKE scenarios. It records submitted and effective worker SKU/shape, node bounds/current count, capacity assumptions, worker-autoscaling thresholds/recovery behavior, price fidelity, and per-field provenance: `human-provided`, `agent-supplied`, `platform-defaulted`, `catalog-sourced`, `engine-inferred`, or `unsupported/fallback`. MCP-created values are `agent-supplied`; do not narrate an agent calibration as a human observation. For a targeted CPU HPA override, send `"autoscalingTargetCpu": 70` at create or update time. `"scaleOutCpuThreshold"`, `"scaleOutCpuPercent"`, and `"autoscaleTargetCpuPercent"` are accepted aliases (the latter two are Grok-compatible). When more than one target name is sent, every value must be the same. CWM retains the caller’s 70% target and fills unrelated autoscaling fields from the provider profile. **These four fields are simulation-wide, top-level only** — the engine applies one `configOverride` identically to every resource's scale decision by default. To make ONE resource scale at a different CPU target than the rest of the simulation (e.g. a GKE cluster scaling out at 60% while an EC2 fleet in the same simulation scales out at 80%), set `characteristics.scaleOutCpuThreshold` and/or `characteristics.scaleInCpuThreshold` on that specific resource instead — this per-resource value wins over the simulation-wide default for that resource's scale decisions only, and every other resource is unaffected. It works the same way whether the resource is a plain compute instance or a Kubernetes cluster. A curated list of near-miss top-level spellings (e.g. `autoscalingTargetCpuPercent`, the real Kubernetes HPA API name `targetCPUUtilizationPercentage`) nested under `characteristics` is rejected with 400 naming the correct field, instead of being silently stripped. Generic Kubernetes has strict scope limits: CWM models worker-node capacity, CPU/ utilization, autoscaling, and worker-resource recovery only. The control-plane management fee is cost-only. It does **not** simulate or report GKE control-plane CPU, master/API throttling, or control-plane cooldown behavior. `metrics.memoryUsage` is synthetic aggregate utilization; it cannot be translated into per-pod GiB, working-set demand, cgroup limits, eviction, OOM timing, or OOMKill behavior. Only an explicit `runtimeMemoryProfile` and its telemetry can support runtime-memory claims. Unknown worker SKUs and generic worker prices are explicitly marked `unsupported/fallback` with estimated fidelity; unknown GPU rates additionally state when their cost is excluded from aggregate totals. **Zero-configuration Kubernetes memory evidence:** A standard literal `apps/v1` Deployment or StatefulSet is enough for GitHub PR simulations to report observed configuration. CWM records the workload kind/name/namespace, source manifest, literal `resources.requests.memory`, literal `resources.limits.memory`, replica configuration, and supported literal CPU HPA settings when present. No CWM manifest, annotation, repository convention, or customer setup is required. Request-only and limit-only manifests are reported as partial evidence; ambiguous multi-container workloads and non-literal values are reported as such without choosing a container. These observations are distinct from runtime evidence: configured memory is capacity, not application working-set demand. Kubernetes/GKE clusters can opt into an explicit per-replica `characteristics.kubernetesMemoryProfile` (limit/request GiB, baseline/burst working set, a `loadBreakpoints` curve keyed on per-replica RPS, source/confidence, `restartDelaySteps` (simulation steps; the deprecated `restartDelaySeconds` input is accepted as a compatibility alias), and an opt-in `memoryAwareAutoscaling` threshold). CWM never infers per-pod memory, OOM likelihood, or working-set demand from `memoryUsage`, CPU, RPS, or the configured limit alone. Each step, `metrics.kubernetesMemory[]` reports every kubernetes resource's per-replica modeled demand vs. limit, headroom, and evaluation status (`not_evaluated` without runtime evidence, `within_limit`/`limit_exceeded` when evaluated). Multi-pool clusters emit one row per pool and evaluate only pools with their own profile. Replicas that cross the limit are attributed an OOMKill event carrying the triggering demand, configured limit, and profile source/confidence, then serve a bounded restart delay measured in simulation steps before rejoining capacity — surviving replicas absorb redistributed traffic through the same capacity-ceiling path used for other scaling events, with latency/error effects visible exactly as other capacity-loss events. CPU-only HPA (the default) never reacts to this signal; `memoryAwareAutoscaling: true` lets the modeled pressure also drive scale-out, and `metrics.kubernetesMemory[].autoscaleDriver` always reports which signal (`cpu`/`memory`/`both`/`none`) drove the decision. Node/replica billing is unaffected — this path only changes capacity, latency, error, and event telemetry. GitHub App PR simulations may optionally attach that profile through the bounded `.cwm/runtime-memory.yaml` (or `.json`) manifest. A Kubernetes section has the following shape: ``` version: 1 iacFiles: [k8s/api.yaml] kubernetes: profiles: api: baselineWorkingSetGiB: 0.5 burstWorkingSetGiB: 0.75 headroomGiB: 0.1 loadBreakpoints: - { perReplicaRps: 100, workingSetGiB: 0.5 } - { perReplicaRps: 600, workingSetGiB: 0.75 } extrapolation: linear source: user-provided confidence: high restartDelaySteps: 30 maxConcurrentRestartFraction: 0.333333 memoryAwareAutoscaling: true memoryScaleOutThresholdRatio: 0.85 ``` The profile map key is an exact Deployment or StatefulSet workload name. `iacFiles` is the complete, bounded source scope: the manifest and every declared file are fetched independently at each immutable base/head ref, so a manifest-only change does not trigger repository-wide discovery. The selected source YAML supplies observed literal `spec.template.spec.containers` memory request and limit; the optional profile supplies only working-set, headroom, restart, source, confidence, and optional memory-HPA behavior. The attachment is rejected without inference when the workload is missing, duplicated, namespace-ambiguous, a ReplicaSet or other unsupported kind, has multiple containers, lacks a literal memory limit, or has malformed/unsupported profile fields. Rejected attachments remain neutral (`not_evaluated`) and identify only the manifest, workload, and safe rejection reason in PR provenance; repository contents and working-set curve values are not copied into public comments or read responses. --- ## Docs & Discovery - [Docs](https://www.cloudworldmodel.ai/docs) - Documentation hub and developer entry point - [Compare](https://www.cloudworldmodel.ai/compare) - Compare cloud simulation approaches - [Glossary](https://www.cloudworldmodel.ai/glossary) - Definitions for cloud simulation terminology - [FAQ](https://www.cloudworldmodel.ai/faq) - Frequently asked questions - [About](https://www.cloudworldmodel.ai/about) - About Cloud World Model - [Pre-provision multi-cloud simulation guide](https://www.cloudworldmodel.ai/guides/pre-provision-multi-cloud-simulation) - Explore multi-cloud designs before provisioning resources - [RL autoscaling environments guide](https://www.cloudworldmodel.ai/guides/rl-autoscaling-environments) - Learn about RL autoscaling simulation environments - [Chaos without production guide](https://www.cloudworldmodel.ai/guides/chaos-without-production) - Explore failure testing in simulation - [llms.txt](https://www.cloudworldmodel.ai/llms.txt) - Concise summary for AI tools (llmstxt.org format) - [OpenAPI spec (JSON)](https://www.cloudworldmodel.ai/api-docs/openapi.json) - Machine-readable OpenAPI 3.1 specification in JSON format (full 90-route spec) — consumable by AI coding assistants and API clients - [OpenAPI spec (YAML)](https://www.cloudworldmodel.ai/openapi.yaml) - Machine-readable OpenAPI 3.1 specification in YAML format (served from project root) - [Swagger UI](https://www.cloudworldmodel.ai/api-docs) - Interactive API explorer powered by Swagger UI - [Sitemap](https://www.cloudworldmodel.ai/sitemap.xml) - XML sitemap of all public pages - [Canvas Cloud AI](https://canvascloud.ai) - Primary product that Cloud World Model complements