awesome-private-inference a registry of TEE-verified inference

What can these providers actually prove?

Every day we ask each confidential-inference provider for an attestation and check what it contains. A row is verified only when every layer that provider's architecture should be able to prove, it proves. Anything we could not reach is marked as such rather than scored — an unreachable endpoint is not a privacy finding. Methodology →

How much to trust this page

Over 169 days and 3372 observations, this instrument has produced 0 verification failures. Every failing observation in its history was a transport error (810) or an invalid response with no recorded reason (224). By the rule of three, zero events in 3372 trials puts the 95% upper bound on the per-observation verification-failure rate at 0.089%. That is the honest reading: not "providers are safe", but "if this check fails, it fails less often than about once in 1124 observations — or it is not wired to anything." We cannot yet distinguish those two cases.

Targets verified
2 / 17
4 partial · 11 unreachable
Builds reviewed
7 / 86
backlog 79 · last 35d ago
Cells carrying signal
27%
124 of 170 matrix cells are dashes
Unmeasurable providers
1
venice expose no stable version identity

Coverage by target

Proven counts the layers this provider's shape is expected to prove — the expectation is set per architecture in REQUIRED_LAYERS_BY_SHAPE, and that bar is an editorial judgment, not a measurement. Different providers have different denominators, so compare the missing column, not the fraction alone.

Provider Model Status Proven Unproven / reason
aci-gateway inference.phala.com verified 6/6 —
aci-gateway tee.redpill.ai verified 6/6 —
aci-gateway api.redpill.ai invalid — no response
near-ai zai-org/GLM-5.1-FP8 partial 6/7 backend_attested
tinfoil router partial 3/5 client_nonce_supported, runtime_config_fully_attested
not shown on the matrix below: client_nonce_supported, code_measurement_reproducible, hpke_pubkey_attested, runtime_config_fully_attested, tls_pubkey_pinned
venice e2ee-gpt-oss-120b-p partial 4/9 code_measurement_reproducible, compose_hash_committed, gpu_attested, prod_os_image, serving_code_attested
not shown on the matrix below: code_measurement_reproducible
venice e2ee-qwen3-6-35b-a3b invalid — no response
not shown on the matrix below: code_measurement_reproducible
venice e2ee-qwen3-6-35b-a3b-uncensored-p partial 5/9 code_measurement_reproducible, compose_hash_committed, prod_os_image, serving_code_attested
not shown on the matrix below: code_measurement_reproducible
chutes DeepSeek-V3.2-TEE unreachable — RuntimeError: HTTP 429 on /instances/6f168c4b-aa9a-48fb-89ef-a1a35f10f779/evidence?nonce=c
chutes GLM-5.1-TEE unreachable — RuntimeError: HTTP 429 on /instances/ca6f001c-85b3-4dd3-bf36-6561997325c6/evidence?nonce=7
chutes GLM-5.2-TEE unreachable — RuntimeError: HTTP 429 on /instances/b0a0a56d-8a88-4993-a601-767e6c980315/evidence?nonce=f
chutes Kimi-K2.6-TEE unreachable — RuntimeError: HTTP 429 on /instances/8f34b201-9043-4b16-9d1b-71da4518663f/evidence?nonce=0
chutes Qwen3-32B-TEE unreachable — RuntimeError: HTTP 429 on /instances/59ca08d8-e6b8-459e-a2ad-2e2ac23b2189/evidence?nonce=d
chutes gemma-4-31B-TEE unreachable — RuntimeError: HTTP 429 on /instances/c4b71bf9-28ce-47b2-8801-bc79611eb9f9/evidence?nonce=c
near-ai openai/gpt-oss-120b unreachable — HTTP 503: {"error":{"message":"Provider error: Model 'openai/gpt-oss-120b' not found. It's
venice e2ee-gemma-4-31b unreachable — no TEE attestation available for this model (404)
not shown on the matrix below: code_measurement_reproducible
venice e2ee-glm-5-1 unreachable — no TEE attestation available for this model (404)
not shown on the matrix below: code_measurement_reproducible

How often does the attested code change?

Any client that pins an enclave measurement is betting the measurement holds still. It does not. Deploys counts transitions to a version string never seen before. Corrected accounts for deploys that begin and end between two probes: with T deploys seen over n−1 daily intervals the rate estimate is −ln(1−T/(n−1))/Δ.

Read revisits only on instance-sampled rows. There the version comes from whichever backend answered, so a return to an earlier value means we reached a different instance — Chutes' 19 revisits are one deploy plus a fleet that never finished draining, not 19 rollouts. On control-plane rows the version is read from a single document (NEAR from the gateway's compose, Tinfoil from the release feed), so it cannot show fleet structure at all and zero revisits there is guaranteed by construction rather than observed. Fleet size is not identifiable from once-a-day sampling on any row.

Target Version source Observations Distinct versions Deploys Revisits Days per deploy Corrected
aci-gateway/api.redpill.ai control-plane 51 11 10 1 5.0 4.5
aci-gateway/inference.phala.com control-plane 51 11 10 0 5.0 4.5
aci-gateway/tee.redpill.ai control-plane 51 11 10 0 5.0 4.5
chutes/DeepSeek-V3.2-TEE instance-sampled 106 2 1 0 110.0 109.5
chutes/GLM-5-TEE instance-sampled 43 2 1 6 43.0 42.5
chutes/GLM-5.1-TEE instance-sampled 47 1 0 0 — —
chutes/GLM-5.2-TEE instance-sampled 46 1 0 0 — —
chutes/Kimi-K2.6-TEE instance-sampled 105 2 1 4 110.0 109.5
chutes/Qwen3-32B-TEE instance-sampled 105 2 1 34 110.0 109.5
chutes/gemma-4-31B-TEE instance-sampled 105 2 1 22 110.0 109.5
near-ai/Qwen/Qwen3-30B-A3B-Instruct-2507 control-plane 32 10 9 0 3.4 2.9
near-ai/openai/gpt-oss-120b control-plane 99 20 19 0 5.3 4.8
near-ai/zai-org/GLM-5-FP8 control-plane 31 10 9 0 3.4 2.9
near-ai/zai-org/GLM-5.1-FP8 control-plane 141 27 26 0 5.4 4.9
redpill/phala/gpt-oss-20b control-plane 33 1 0 0 — —
redpill/phala/qwen-2.5-7b-instruct control-plane 34 1 0 0 — —
tinfoil/gemma4-31b control-plane 87 7 6 0 14.5 14.0
tinfoil/gpt-oss-120b control-plane 82 2 1 0 82.0 81.5
tinfoil/llama3-3-70b control-plane 93 3 2 0 46.5 46.0
tinfoil/router control-plane 164 44 43 0 3.8 3.3

Per-layer detail

Cells: ✅ verified, ❌ rejected, ○ awaiting our review, — required but not exposed, — not applicable to this architecture. This matrix shows 10 of the 15 layers the probe records; the remainder are listed per row in the coverage table above.

Provider Model Shape Nonce bound TDX quote report_data binds key GPU attested Key derives to addr compose_hash committed Prod OS image Serving code attested Backend attested Attested serving forced lat
aci-gateway inference.phala.com aci-gateway ✅ ✅ ✅ — — ✅ ✅ — ❌ ✅ 0.0s
aci-gateway tee.redpill.ai aci-gateway ✅ ✅ ✅ — — ✅ ✅ — ❌ ✅ 0.0s
aci-gateway api.redpill.ai aci-gateway ✅ ✅ ✅ — — ✅ ✅ — ❌ ❌ 0.0s
near-ai zai-org/GLM-5.1-FP8 tdx+gpu ✅ ✅ ✅ ✅ ✅ ✅ — — ❌ — 2.7s
tinfoil router tinfoil-sev-snp-v2 — — — — — — — — — — 0.56s
venice e2ee-gpt-oss-120b-p venice ✅ ✅ ✅ — ✅ — — — ❌ — 1.8s
venice e2ee-qwen3-6-35b-a3b venice ✅ ❌ ✅ ✅ ✅ ✅ — — ❌ — 2.09s
venice e2ee-qwen3-6-35b-a3b-uncensored-p venice ✅ ✅ ✅ ✅ ✅ — — — ❌ — 2.62s
chutes DeepSeek-V3.2-TEE chutes-tee — — — — — — — — — — 0.53s
chutes GLM-5.1-TEE chutes-tee — — — — — — — — — — 0.69s
chutes GLM-5.2-TEE chutes-tee — — — — — — — — — — 0.75s
chutes Kimi-K2.6-TEE chutes-tee — — — — — — — — — — 0.51s
chutes Qwen3-32B-TEE chutes-tee — — — — — — — — — — 0.69s
chutes gemma-4-31B-TEE chutes-tee — — — — — — — — — — 0.71s
near-ai openai/gpt-oss-120b tdx+gpu — — — — — — — — — — 0.28s
venice e2ee-gemma-4-31b venice — — — — — — — — — — 0.12s
venice e2ee-glm-5-1 venice — — — — — — — — — — 0.18s

Editorial notes

Hand-written, not computed, and therefore able to go stale. Each carries the date it was last checked against the data. Anything not on this list that you see above is machine-derived.

How this page is generated

Daily at 13:17 UTC, probe-daily.yml runs python -m probes.collect, writes data/snapshots/YYYY-MM-DD.json, recomputes data/quality.json (the self-check driving the four tiles above), re-renders this page, and commits. Reproduce with run a probe locally; grade this page with python -m probes.quality.