Cloud World Model on Smithery

    Simulation Accuracy Benchmark

    This page publishes a reproducible, company-owned cwm-bench measurement of the canonical AWS architecture and clearly separates it from engine predictions and provider-documentation references.

    See also: Simulation Fidelity — benchmark data and accuracy ranges for all five cloud providers.

    AWS scored app CPU, in-VPC internal-load-balancer latency, goodput and CRUD errors come from the pinned owned cwm-bench campaign. Cost uses the us-east-2 price list. Tuning uses 10 / 100 / 500 RPS; 1,000 RPS, later-day and us-west-2 are holdouts. See the Methodology section below for full source citations.

    GCP, Azure, OCI, and DigitalOcean scores use documentation references, not owned measurements. They are shown separately from canonical AWS and are not a cross-provider ranking.

    OpenShift reference scenarios
    Reference-only
    OpenShift is a platform overlay, not a sixth cloud provider, so it is not included in the measured provider scorecard yet.
    Cost: Estimated
    Latency, CPU, throughput, errors: Extrapolated

    Independently sourced coverage currently includes rosa-hcp, rosa-classic, aro, openshift-dedicated, self-managed across 6 region-specific scenarios. Platform fees are estimated from published product/pricing information; performance behavior is not a topology-matched public load test.

    View source-backed OpenShift reference output
    All Providers at a Glance
    Scores are grouped by evidence basis, not ranked together. Documentation references have not been checked against owned measurements. Click a row to jump to its detailed results below.
    ProviderScoring basisOverall
    Owned lean-weight measurement comparison (cost: public price list); typical holdout shown separately below
    AWS (lean)
    Loading basis…
    Documentation-reference comparisons — not yet measured
    GCP
    Loading basis…
    Azure
    Loading basis…
    Oracle Cloud (OCI)
    Loading basis…
    DigitalOcean
    Loading basis…

    Click any row to load that provider's full benchmark results below.

    CPU Accuracy by Scenario
    Simulated vs. reference CPU utilization at each traffic scenario. CPU is one of the largest single drivers of each provider's score, so this shows why a provider lands where it does — and makes future calibration changes visible at a glance. Each cell colors the simulated value by how closely it tracks the reference (green ≤ 10%, yellow ≤ 25%, red above 25% off).
    Provider / scoring basis
    Idle
    Normal
    Peak
    Burst
    Owned measurement comparison
    AWSLoading basis…
    Documentation references — not checked against owned measurements
    GCPLoading basis…
    AzureLoading basis…
    Oracle Cloud (OCI)Loading basis…
    DigitalOceanLoading basis…

    AWS cells show predicted application-server CPU against measured application-server CPU; other providers show simulated output against documentation-sourced references. Click any row to load that provider's full benchmark below.

    6th-Gen AWS Instance Accuracy
    Per-instance benchmark scores for six 6th-generation AWS instance types (m6i, c6i, r6i) across all four traffic scenarios. Scores are derived from the same reference scenarios used by the canonical AWS benchmark above.
    InstanceOverall

    Each row benchmarks the simulator against doc-sourced reference values for that specific instance type. Cost, Latency, and Perf are category averages across the four traffic scenarios.

    Select a provider to load its canonical benchmark automatically.

    AWS canonical comparison scores owned cwm-bench app-host CPU and k6 in-VPC internal-LB latency against seeded engine predictions, with owned goodput/errors and AWS price-list cost. Idle/Normal/Peak are fit rungs; Burst is a holdout. Later-day and second-region have no separate predictions.
    Tune the Architecture
    Choose a cloud provider then adjust instance counts and types to see how simulation accuracy changes for your setup. Reference values stay fixed at the canonical AWS architecture — you're exploring how the simulator responds across providers and resource sizes.
    2
    110

    Results update the architecture diagram and scenario tables below.

    Canonical Architecture

    Traffic Scenario Comparisons

    Each table shows Reference vs. Simulated values, the percentage delta, and an accuracy badge (green ≥ 90%, yellow ≥ 75%, red below 75%).

    Pricing accuracy is validated continuously via automated drift checks in CI. See the Simulation Fidelity page for per-provider cost benchmark details across all five cloud providers.

    Want to benchmark a different architecture? Open the Workspace to build and simulate any topology.

    Try the Simulation Yourself

    Load the same 3-tier AWS scenario in the interactive workspace and compare what you see with the reference values on this page.