What can these providers actually prove?
Every day we ask each confidential-inference provider for an attestation and check what it contains. A row is verified only when every layer that provider's architecture should be able to prove, it proves. Anything we could not reach is marked as such rather than scored — an unreachable endpoint is not a privacy finding. Methodology →
How much to trust this page
Over 132 days and 2743 observations, this instrument has produced 0 verification failures. Every failing observation in its history was a transport error (671) or an invalid response with no recorded reason (157). By the rule of three, zero events in 2743 trials puts the 95% upper bound on the per-observation verification-failure rate at 0.109%. That is the honest reading: not "providers are safe", but "if this check fails, it fails less often than about once in 914 observations — or it is not wired to anything." We cannot yet distinguish those two cases.
Coverage by target
Proven counts the layers this provider's shape is expected to prove — the
expectation is set per architecture in
REQUIRED_LAYERS_BY_SHAPE, and that bar is an editorial judgment, not a
measurement. Different providers have different denominators, so compare the
missing column, not the fraction alone.
| Provider | Model | Status | Proven | Unproven / reason |
|---|---|---|---|---|
| aci-gateway | inference.phala.com | verified | 6/6 | — |
| aci-gateway | tee.redpill.ai | verified | 6/6 | — |
| aci-gateway | api.redpill.ai | invalid | — | no response |
| near-ai | zai-org/GLM-5.1-FP8 | partial | 6/7 | backend_attested |
| tinfoil | router | partial | 3/5 |
client_nonce_supported, runtime_config_fully_attested
not shown on the matrix below:
client_nonce_supported, code_measurement_reproducible, hpke_pubkey_attested, runtime_config_fully_attested, tls_pubkey_pinned
|
| venice | e2ee-glm-5-1 | partial | 6/9 |
code_measurement_reproducible, prod_os_image, serving_code_attested
not shown on the matrix below:
code_measurement_reproducible
|
| venice | e2ee-qwen3-6-35b-a3b | partial | 6/9 |
code_measurement_reproducible, prod_os_image, serving_code_attested
not shown on the matrix below:
code_measurement_reproducible
|
| venice | e2ee-qwen3-6-35b-a3b-uncensored-p | partial | 5/9 |
code_measurement_reproducible, compose_hash_committed, prod_os_image, serving_code_attested
not shown on the matrix below:
code_measurement_reproducible
|
| chutes | DeepSeek-V3.2-TEE | unreachable | — | RuntimeError: HTTP 429 on /instances/38f7cf24-f141-426c-ac3c-e8b4f33f9acf/evidence?nonce=c |
| chutes | GLM-5.1-TEE | unreachable | — | RuntimeError: HTTP 429 on /instances/bec3b3e2-a641-48be-8b2f-eeef4e8c1607/evidence?nonce=b |
| chutes | GLM-5.2-TEE | unreachable | — | RuntimeError: HTTP 429 on /instances/5bc1b776-b167-40d8-9585-d6acd5297b47/evidence?nonce=f |
| chutes | Kimi-K2.6-TEE | unreachable | — | RuntimeError: HTTP 429 on /instances/b89c6c42-fb04-406b-a907-269ad4f22e08/evidence?nonce=e |
| chutes | Qwen3-32B-TEE | unreachable | — | RuntimeError: HTTP 429 on /instances/ff977ee9-b4f8-4e03-b16d-b13176079196/evidence?nonce=c |
| chutes | gemma-4-31B-TEE | unreachable | — | RuntimeError: HTTP 429 on /instances/850578c8-3a69-4e77-82ab-9353cbe0e110/evidence?nonce=3 |
| near-ai | openai/gpt-oss-120b | unreachable | — | HTTP 503: {"error":{"message":"Provider error: Model 'openai/gpt-oss-120b' not found. It's |
| venice | e2ee-gemma-4-31b | unreachable | — |
HTTP 502: {"error":"TEE attestation request failed. The Trusted Execution Environment prov
not shown on the matrix below:
code_measurement_reproducible
|
| venice | e2ee-gpt-oss-120b-p | unreachable | — |
HTTP 502: {"api_version":"aci/1","workload_keyset_digest":"sha256:3e8c94d2204afbc999e63ccb
not shown on the matrix below:
code_measurement_reproducible
|
How often does the attested code change?
Any client that pins an enclave measurement is betting the measurement holds still. It does not. Deploys counts transitions to a version string never seen before. Corrected accounts for deploys that begin and end between two probes: with T deploys seen over n−1 daily intervals the rate estimate is −ln(1−T/(n−1))/Δ.
Read revisits only on instance-sampled rows. There the version
comes from whichever backend answered, so a return to an earlier value means we reached a
different instance — Chutes' 19 revisits are one deploy plus a fleet that never finished
draining, not 19 rollouts. On control-plane rows the version is read from a
single document (NEAR from the gateway's compose, Tinfoil from the release feed), so it
cannot show fleet structure at all and zero revisits there is guaranteed by construction
rather than observed. Fleet size is not identifiable from once-a-day sampling on any row.
| Target | Version source | Observations | Distinct versions | Deploys | Revisits | Days per deploy | Corrected |
|---|---|---|---|---|---|---|---|
| aci-gateway/api.redpill.ai | control-plane | 14 | 5 | 4 | 0 | 3.2 | 2.7 |
| aci-gateway/inference.phala.com | control-plane | 14 | 5 | 4 | 0 | 3.2 | 2.7 |
| aci-gateway/tee.redpill.ai | control-plane | 14 | 5 | 4 | 0 | 3.2 | 2.7 |
| chutes/DeepSeek-V3.2-TEE | instance-sampled | 73 | 2 | 1 | 0 | 73.0 | 72.5 |
| chutes/GLM-5-TEE | instance-sampled | 43 | 2 | 1 | 6 | 43.0 | 42.5 |
| chutes/GLM-5.1-TEE | instance-sampled | 13 | 1 | 0 | 0 | — | — |
| chutes/GLM-5.2-TEE | instance-sampled | 13 | 1 | 0 | 0 | — | — |
| chutes/Kimi-K2.6-TEE | instance-sampled | 72 | 2 | 1 | 4 | 73.0 | 72.5 |
| chutes/Qwen3-32B-TEE | instance-sampled | 73 | 2 | 1 | 24 | 73.0 | 72.5 |
| chutes/gemma-4-31B-TEE | instance-sampled | 73 | 2 | 1 | 18 | 73.0 | 72.5 |
| near-ai/Qwen/Qwen3-30B-A3B-Instruct-2507 | control-plane | 32 | 10 | 9 | 0 | 3.4 | 2.9 |
| near-ai/openai/gpt-oss-120b | control-plane | 99 | 20 | 19 | 0 | 5.3 | 4.8 |
| near-ai/zai-org/GLM-5-FP8 | control-plane | 31 | 10 | 9 | 0 | 3.4 | 2.9 |
| near-ai/zai-org/GLM-5.1-FP8 | control-plane | 104 | 19 | 18 | 0 | 5.8 | 5.3 |
| redpill/phala/gpt-oss-20b | control-plane | 33 | 1 | 0 | 0 | — | — |
| redpill/phala/qwen-2.5-7b-instruct | control-plane | 34 | 1 | 0 | 0 | — | — |
| tinfoil/gemma4-31b | control-plane | 87 | 7 | 6 | 0 | 14.5 | 14.0 |
| tinfoil/gpt-oss-120b | control-plane | 82 | 2 | 1 | 0 | 82.0 | 81.5 |
| tinfoil/llama3-3-70b | control-plane | 93 | 3 | 2 | 0 | 46.5 | 46.0 |
| tinfoil/router | control-plane | 127 | 33 | 32 | 0 | 4.0 | 3.4 |
Per-layer detail
Cells: ✅ verified, ❌ rejected, ○ awaiting our review, — required but not exposed, — not applicable to this architecture. This matrix shows 10 of the 15 layers the probe records; the remainder are listed per row in the coverage table above.
| Provider | Model | Shape | Nonce bound | TDX quote | report_data binds key | GPU attested | Key derives to addr | compose_hash committed | Prod OS image | Serving code attested | Backend attested | Attested serving forced | lat |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| aci-gateway | inference.phala.com | aci-gateway | ✅ | ✅ | ✅ | — | — | ✅ | ✅ | — | ❌ | ✅ | 0.0s |
| aci-gateway | tee.redpill.ai | aci-gateway | ✅ | ✅ | ✅ | — | — | ✅ | ✅ | — | ❌ | ✅ | 0.0s |
| aci-gateway | api.redpill.ai | aci-gateway | ✅ | ✅ | ✅ | — | — | ✅ | ❌ | — | ❌ | ❌ | 0.0s |
| near-ai | zai-org/GLM-5.1-FP8 | tdx+gpu | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | — | — | ○ | — | 1.89s |
| tinfoil | router | tinfoil-sev-snp-v2 | — | — | — | — | — | — | — | — | — | — | 0.44s |
| venice | e2ee-glm-5-1 | venice | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | — | — | ❌ | — | 3.74s |
| venice | e2ee-qwen3-6-35b-a3b | venice | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | — | — | ❌ | — | 3.07s |
| venice | e2ee-qwen3-6-35b-a3b-uncensored-p | venice | ✅ | ✅ | ✅ | ✅ | ✅ | — | — | — | ❌ | — | 2.97s |
| chutes | DeepSeek-V3.2-TEE | chutes-tee | — | — | — | — | — | — | — | — | — | — | 0.57s |
| chutes | GLM-5.1-TEE | chutes-tee | — | — | — | — | — | — | — | — | — | — | 4.36s |
| chutes | GLM-5.2-TEE | chutes-tee | — | — | — | — | — | — | — | — | — | — | 4.12s |
| chutes | Kimi-K2.6-TEE | chutes-tee | — | — | — | — | — | — | — | — | — | — | 0.81s |
| chutes | Qwen3-32B-TEE | chutes-tee | — | — | — | — | — | — | — | — | — | — | 4.35s |
| chutes | gemma-4-31B-TEE | chutes-tee | — | — | — | — | — | — | — | — | — | — | 4.08s |
| near-ai | openai/gpt-oss-120b | tdx+gpu | — | — | — | — | — | — | — | — | — | — | 0.2s |
| venice | e2ee-gemma-4-31b | venice | — | — | — | — | — | — | — | — | — | — | 17.68s |
| venice | e2ee-gpt-oss-120b-p | venice | — | — | — | — | — | — | — | — | — | — | 0.65s |
Editorial notes
Hand-written, not computed, and therefore able to go stale. Each carries the date it was last checked against the data. Anything not on this list that you see above is machine-derived.
- 2026-08-18 Current RedPill uses ACI. The live tee.redpill.ai gateway scores Stage 0 even though all six ACI protocol checks now pass, including the strict production-OS appraisal — the deployment moved from the dstack dev image to prod 0.5.9 on or before 2026-08-18. Receipts verify end to end. Stage 0 rests on the legacy /v1/attestation/report, which ignores its model parameter and attests any name; public logs with raw upstream error detail; and an operator root-key input. api.redpill.ai shares the same attested keyset but is not TEE-only. A session audit accepted 163 of 237 records, rejecting only Chutes for missing evidence.
- 2026-06-18 Chutes' serving code is not measured. serve.py on the prompt-plaintext path is CFSV-excluded and in no RTMR, and the model name is not bound to the quote. A passing quote proves genuine TDX running a Chutes base image, not which model on which code.
- 2026-05-09 NEAR's gateway gap depends on the client. ALLOWED_COMPOSE_HASHES is unset server-side, so the gateway alone does not pin code. A closed-chain client that checks compose_hash against the on-chain set on Base closes it; a client that trusts the gateway does not.
- 2026-08-12 Some provider-adapter JWTs are decoded without signature checks. Phala's private-ai-verifier still passes verify_signature=False on NVIDIA and Intel Trust Authority tokens. This affects adapter conclusions derived from those tokens. The current ACI client verifies its gateway TDX quote through native DCAP.
- 2026-08-10 The bar is hand-set, and one entry was wrong. REQUIRED_LAYERS_BY_SHAPE is a hand-edited dict with no changelog. Venice's set excluded the two layers where its prompt-path exposure actually lives, so Venice scored a full row until an outside audit pointed at them; corrected 2026-08-10. Denominators still differ per architecture, so compare the unproven layers rather than the fractions, and treat the bar as editorial until it is derived from provider claims (issue #6).
How this page is generated
Daily at 13:17 UTC,
probe-daily.yml
runs python -m probes.collect, writes data/snapshots/YYYY-MM-DD.json,
recomputes data/quality.json (the self-check driving the four tiles above),
re-renders this page, and commits. Reproduce with
run a probe locally;
grade this page with python -m probes.quality.