Last month a single AWS Graviton5 machine ran our 22-query TPC-H 1 TB benchmark in 80 seconds, and I closed that post with a promise: the horse race between the clouds' Arm chips would continue. So we held a proper race. Six GizmoSQL clusters on AWS, Google Cloud and Azure, all 64 vCPUs with local NVMe. Each ran the full 1 TB benchmark twice over: once from local NVMe and once from a DuckLake catalog in that cloud's own object storage, with both the current DuckDB and the DuckDB 2.0 alpha.
That is twelve results, and every one of the 792 query runs succeeded. The whole race, from the first cluster created to the last one deleted, took 38 minutes and cost about $37 on GizmoData Cloud. The fastest lap was AWS Graviton5 on local NVMe with the DuckDB 2.0 alpha: 79 seconds for all 22 queries over 1 TB, for $0.30. That beats last month's 80.5 seconds with a third fewer cores.
The setup most teams never get to
Before the horses, the paddock. A fair cross-cloud benchmark is mostly plumbing. Before a single query runs, you need the same engine, the same data and the same settings in three clouds, and you need to tear it all down before it bills you for the weekend.
Doing it yourself means:
- Three cloud accounts, each with billing set up and its 64-vCPU quota raised (Google Cloud also has a separate per-family quota for its Axion machines).
- Networks, Kubernetes clusters or VMs, and local NVMe disks prepared in each cloud.
- TLS certificates, DNS records and load balancers for every server.
- Generating 1 TB of TPC-H data and copying it into three object stores.
- A Postgres catalog per region for DuckLake, plus storage credentials for each cloud.
- Installing and configuring the engine twelve ways (three clouds, two versions, two storage modes).
- Remembering to delete all of it afterwards.
On GizmoData Cloud it was:
- One API call per cluster, naming the cloud region, the machine size and the GizmoSQL release channel.
- A startup SQL script that loads the 1 TB TPC-H database onto the cluster's local NVMe as it boots.
- A DuckLake catalog of the same data, in that cloud's own object storage, mounted at creation.
- A load balancer, TLS certificate and hostname for every cluster, on all three clouds, with nothing to configure.
- A "ready to connect" signal, so the benchmark starts the moment the data is there.
- Auto-terminate as a safety net, and a delete call as soon as each racer finishes.
The AWS clusters were up with the full terabyte loaded on NVMe and verified about 5 minutes after the API call, and the Google Cloud clusters about 7. Every racer was deleted the moment its last query finished, which is why the whole race cost about $37.
The results
Each number is the sum of the mean times of the 22 TPC-H queries (lower is faster). Under it is what that run costs on GizmoData Cloud, Enterprise edition. Stable is GizmoSQL v1.41.0 with DuckDB 1.5.6; edge is GizmoSQL v1.41.0 with the DuckDB 2.0 alpha.
| 64 vCPUs | NVMe · stable | NVMe · edge | DuckLake · stable | DuckLake · edge |
|---|---|---|---|---|
AWSr9gd.16xlarge (Graviton5) | 108 s $0.41 | 79 s $0.30 | 561 s $2.16 | 370 s $1.42 |
Google Cloudc4a-highmem-64-lssd (Axion) | 140 s $0.43 | 97 s $0.30 | 428 s $1.32 | 258 s $0.80 |
AzureStandard_E64pds_v6 (Cobalt 100) | 171 s $0.53 | 134 s $0.41 | 398 s $1.23 | 314 s $0.96 |
And the gold medal goes to…
We scored six events, each cloud on its best channel (edge).
- Google Cloud takes four golds: the fastest DuckLake run, the cheapest run on both local NVMe and DuckLake, and the biggest DuckDB 2.0 speed-up.
- AWS takes gold for raw speed: the fastest local NVMe run of the race, 79 seconds for all 22 queries.
- Azure takes gold for the smallest object-storage penalty: its DuckLake run took only 2.3 times as long as its NVMe run.
The cheapest NVMe run was a photo finish: Google Cloud beat AWS by a quarter of a cent.
And the engine itself wins in every lane: the edge channel, with the DuckDB 2.0 alpha, beat the stable channel in all six combinations of cloud and storage.
What we learned
DuckDB 2.0 is a big step, and it helps most on object storage
The edge channel cut total query time by 22% to 31% on local NVMe and by 21% to 40% on DuckLake, on the same hardware with the same data. That is a large gain from an engine upgrade alone, and on GizmoData Cloud it is one field in the create call.
Credit for this one goes to the DuckDB team: we changed nothing but the engine version. The likely main cause is the asynchronous I/O that now runs through DuckDB 2.0. Its I/O layer scales independently of query processing, so far more reads from object storage are in flight at once. That is exactly the DuckLake case, where every Parquet file comes over the network, and it fits what we saw on AWS, whose cold DuckLake run dropped from 724 to 464 seconds. The DuckDB team says local storage benefits only a little from asynchronous I/O, so the NVMe gain more likely comes from the rest of 2.0, such as the new storage format and wider row-group pruning. We haven't profiled it query by query, so treat that split as our best reading, not a measurement.
Local NVMe: AWS Graviton5 is fastest, and the cheapest run is a tie
On local NVMe, the AWS r9gd.16xlarge ran the full 1 TB benchmark in 79 seconds on the edge channel. On cost, AWS and Google Cloud tie at about $0.30 per full pass: Google Cloud's lower hourly price makes up for its longer run (97 seconds).
DuckLake on object storage: the order flips
Reading Parquet from object storage, Google Cloud ran the fastest pass of the race: 258 seconds on the edge channel, against 314 on Azure and 370 on AWS. On the stable channel Azure came in just ahead of Google Cloud (398 against 428 seconds), with AWS at 561. AWS lost most of its time on the first, cold run, when every file comes from S3.
What a 1 TB run costs
A full 22-query pass over 1 TB costs between 30 cents and $2.16 on GizmoData Cloud, depending on the cloud, the storage and the engine version. The newest engine cuts the cost in the same proportion as it cuts the time.
Local NVMe or DuckLake?
Local NVMe was faster on every cloud. A DuckLake pass took 2.3 times as long as an NVMe pass on Azure, 2.7 to 3.1 times on Google Cloud and 4.7 to 5.2 times on AWS. Most of AWS's gap is in the first, cold run, when every Parquet file comes from S3. Azure's gap is the smallest partly because its NVMe run was the slowest of the three; on DuckLake, Azure's edge run (314 seconds) beat AWS (370).
So why use DuckLake at all? NVMe speed comes with conditions: the data has to be loaded onto each cluster when it starts, and it lives only as long as that cluster. A DuckLake catalog keeps the data in your own object storage, where any number of GizmoSQL clusters can read it at once with no load step, and where it outlives every cluster you create or delete. GizmoData Cloud gives you both: mount a DuckLake catalog for shared, durable data, and load a hot working set onto local NVMe with a startup script when you need the speed.
Query by query
Q9, Q13, Q18 and Q3 make up the largest share of every total. On local NVMe with the edge channel, AWS was fastest on 19 of the 22 queries and Google Cloud on the other three (Q16, Q18 and Q20). Azure's gap was widest on Q4, Q8, Q2 and Q5, where it took three to four times as long as AWS.
How we ran it
- Hardware: 64 vCPUs per cluster with local NVMe: AWS
r9gd.16xlarge(Graviton5) in US East 2, Google Cloudc4a-highmem-64-lssd(Axion) in US East 4, AzureStandard_E64pds_v6(Cobalt 100) in East US 2. - Engine: GizmoSQL v1.41.0 on the stable (DuckDB 1.5.6) and edge (DuckDB 2.0 alpha) channels, with the same DuckDB thread setting on every cluster.
- Data: TPC-H at scale factor 1,000. Local NVMe: a DuckDB database on the cluster's NVMe disk, loaded at startup. DuckLake: Parquet in the same region's object storage (S3, Google Cloud Storage, Azure Blob Storage) with a Postgres catalog.
- Runs: each of the 22 queries three times on NVMe and twice on DuckLake. We report the sum of the per-query means, with the cold (first) and warm runs split out for DuckLake.
- Client: our open-source benchmark-flight-sql over Arrow Flight SQL with TLS, from one machine. TPC-H answers are small, so network time is a negligible part of each query.
- Correctness: every one of the 792 query runs returned the same row counts in all 12 cells.
- Cost: run time multiplied by the cluster's GizmoData Cloud hourly price, Enterprise edition (AWS $13.84, Google Cloud $11.13, Azure $11.07 per hour). The Core edition costs less.
One Azure footnote: during this run the Azure pods asked Kubernetes for 62 CPUs instead of 63, to work around a request size that did not fit the node, which we have since corrected. The engine used the same thread count as on the other clouds. And the honest caveat from last time applies: each cell is one benchmark session on one afternoon (October 8, 2026), so a few percent of any gap is noise.
The race isn't over
Our Azure horse is Cobalt 100, Microsoft's first-generation Arm chip, and it still took a gold and beat AWS on DuckLake. Cobalt 200 is next: Microsoft claims more than 50% higher performance than Cobalt 100, and when it reaches Azure's local-NVMe shapes it will very likely change this race again. We will run it the week it does. DuckDB 2.0 is still an alpha, and there are bigger Graviton5 and Axion sizes waiting too.
Which brings me to the real result. The true winner of this race is our customer. Amazon, Microsoft and Google are each spending billions to make their chips beat the other two, and DuckDB gets faster with every release. On GizmoData Cloud you get all of it: a new chip shows up as a new size, and a new DuckDB shows up as a new release channel. Taking the upgrade means changing one string in your create call, not re-platforming. When the order of the horses changes, you move with it. Today, on AWS, that upgrade alone took 27% off the run time and the bill.
Run your own race
Try GizmoData Cloud free
Sign up and enter LAUNCH25 as your Signup Invite Code when you set up your organization: the first 20 sign-ups get $25 of credit, no credit card. Every cluster can mount the same public 1 TB TPC-H dataset we raced on, in its own region, so your first benchmark is one query away.
Run it yourself
Every GizmoData Cloud region carries a public, read-only TPC-H catalog, with one schema per scale factor (sf1, sf10, sf100 and sf1000), so you can repeat the DuckLake half of this race without loading anything. You need a GizmoData Cloud account, an API key from the avatar menu's Profile & API keys, and curl, jq and openssl on your machine. The easiest way to get your project environment ID is the Copy as curl button at the end of the portal's create-cluster wizard. A 64-vCPU cluster needs a plan whose vCPU quota covers it (Business runs up to 128 vCPUs at once); on the trial credit, pick a smaller size such as r9gd.2xlarge and a smaller schema.
Show the API call that creates a racer (curl)
export GIZMODATA_API_KEY=... # Profile & API keys in the portal
export GIZMODATA_PROJECT_ENVIRONMENT_ID=... # from the create wizard's "Copy as curl"
# Letters and digits only: a password containing ':' can't be sent in Basic auth.
GIZMOSQL_PASSWORD=$(openssl rand -hex 16)
CLUSTER_ID=$(curl -sS --fail-with-body -X POST "https://app.gizmodata.com/api/database-clusters/?wait=false" \
-H "Authorization: Bearer $GIZMODATA_API_KEY" \
-H "Content-Type: application/json" \
--data @- <<EOF | jq -r .id
{
"name": "tpch-racer",
"label": "TPC-H racer",
"project_environment_id": "$GIZMODATA_PROJECT_ENVIRONMENT_ID",
"cloud_region": {"label": "aws-us-east-2"},
"compute_instance_size": {"name": "r9gd.16xlarge"},
"enable_local_nvme_ssd": true,
"database_engine": {"name": "duckdb"},
"database_image": {"name": "gizmosql-edge-v1.41.0-slim"},
"license_edition": "enterprise",
"admin_username": "racer",
"admin_password": "$GIZMOSQL_PASSWORD",
"allow_any_ip": true,
"auto_terminate_after_minutes": 120
}
EOF
)
# Wait until the engine answers through its endpoint, then print the connection URI.
until curl -sS "https://app.gizmodata.com/api/database-clusters/$CLUSTER_ID" \
-H "Authorization: Bearer $GIZMODATA_API_KEY" | jq -e .ready_to_connect > /dev/null; do
sleep 15
done
curl -sS "https://app.gizmodata.com/api/database-clusters/$CLUSTER_ID" \
-H "Authorization: Bearer $GIZMODATA_API_KEY" | jq -r '"grpc+tls://\(.ingress_host):443"'
# When you're done:
# curl -sS -X DELETE "https://app.gizmodata.com/api/database-clusters/$CLUSTER_ID" \
# -H "Authorization: Bearer $GIZMODATA_API_KEY"For the other racers, use gcp-us-east4 with c4a-highmem-64-lssd, or azure-eastus2 with Standard_E64pds_v6; for the stable channel, use gizmosql-v1.41.0-slim. auto_terminate_after_minutes is a safety net that deletes the cluster after two hours even if you forget. Then point the benchmark runner at the cluster and the public catalog:
git clone https://github.com/gizmodata/benchmark-flight-sql
cd benchmark-flight-sql
pip install .
benchmark-flight-sql \
--hostname <the host printed above> \
--port 443 \
--username racer \
--password "$GIZMOSQL_PASSWORD" \
--database gizmodata_tpch_aws_us_east_2 \
--schema sf1000 \
--num-query-runs 2The point
Picking a cloud for analytics shouldn't take a month of plumbing, and it shouldn't be a decision you are stuck with. On GizmoData Cloud the same GizmoSQL runs in AWS, Google Cloud and Azure behind one API, so you can measure for yourself which one is fastest or cheapest for your data, and move when the answer changes. It will change: Cobalt 200 and DuckDB 2.0 are both on the way.
If your team is paying a warehouse by the credit, the DBU or the byte scanned for queries like these, this is what 64 vCPUs and a terabyte look like for 30 cents. Start free on GizmoData Cloud, or talk to us about moving a workload.