Research
How Verifiable Inference Works: zkML, Hardware Attestation, and Committed Configuration
September 2, 2026

The first article in this series described the inference audit gap: when an AI system makes a costly decision, the only evidence of which model produced it is the operator's word. Closing the gap means being able to demonstrate two distinct things. First, that the computation was executed correctly, so the output really is what the model computes for that input. Second, that the computation ran in the claimed environment, on the claimed model, under the claimed policy. Different cryptographic tools answer each question, and the honest summary is that both now work, with real published results, at costs that suit audits rather than every token of every session.

Proving the Math: Zero Knowledge Machine Learning

A zero knowledge succinct proof, or SNARK, lets a prover who ran a computation produce a short certificate that a verifier can check without redoing the work and without seeing the internals. zkML applies this to neural networks by compiling the model's arithmetic into a circuit the proof system can execute. EZKL, an open-source toolchain, takes models exported in the standard ONNX format and produces zk-SNARK circuits, so an inference becomes a statement anyone can verify: this output is the correct result of this model on this input.

The results have moved from demo to published research. zkLLM, by Haochen Sun, Jason Li, and Hongyang Zhang of the University of Waterloo, appeared at ACM CCS 2024 and is the first specialized zero knowledge proof system for large language models. It verifies correctness of inference while keeping the model parameters confidential, which matters to providers who treat weights as trade secrets. Using a GPU implementation, it generates a proof of the entire inference process in under 15 minutes for a 13-billion-parameter model, with proofs under 200 kB. Earlier pilots pointed the same direction: Modulus Labs demonstrated verified inference for an on-chain trading model in 2023, and Giza applied zk proof techniques to model deployment on Starknet.

The overhead remains the constraint. Proving a large model's execution costs orders of magnitude more than running it. That shapes the realistic deployment pattern. Verifying a proof is cheap, so high-volume flows can commit a hash of every decision and attach zk proofs to a sampled subset, while high-stakes flows, a payment authorization, a regulated determination, a contractual commitment, can carry a proof for every decision. Audit-grade coverage does not require proving every token of every chat session.

Proving the Environment: Hardware Attestation

The second question is not about arithmetic but about identity: did this really run on the claimed model, in a configuration nobody tampered with? Trusted execution environments answer it with hardware. Intel's SGX pioneered CPU enclaves; server platforms have since moved to Intel TDX and AMD's SEV-SNP, which let a cloud operator run customer workloads while the processor, not the operator, holds the keys to memory. NVIDIA brought the same model to data-center GPUs: H100-class hardware supports a confidential-computing mode in which the GPU itself produces a signed attestation of its firmware and configuration.

The pattern is called remote attestation. The environment measures its own boot chain and configuration, the hardware signs the measurement, and a remote verifier checks the signature against the manufacturer's public key hierarchy before trusting the machine. The strengths are performance, near-native execution, and maturity in cloud fleets. The limits are equally clear. Attestation proves the environment was intact, not that the computation's logic is correct, and it requires trust in the hardware vendor's keys and implementation. A TEE tells you who ran the code; it does not by itself tell you the code did what the model card claims.

Binding the Configuration

Between the math proof and the hardware quote sits a third, simpler tool that is often skipped: a cryptographic commitment to the configuration itself. The model fingerprint is a hash of the weights. The tokenizer and preprocessing pipeline hash to their own values. The policy, the system prompt, the guardrails, the sampling parameters, is a small structured object that hashes like anything else. Concatenate them and you have a configuration fingerprint: one string that identifies exactly which model, under exactly which instructions, produced a response.

Attaching that fingerprint to each response, or to a Merkle root covering a batch of responses, turns "which model answered?" from a documentation question into a lookup. An auditor holding the record and the artifacts can recompute the hashes and confirm the binding. zkML proves the computation was carried out faithfully; hardware attestation proves the environment was what it claimed; the committed configuration states what was claimed in the first place. The three together cover both halves of the problem the first article described.

What a Proof Does Not Tell You

Honesty about limits keeps this technology credible. A correctness proof of a bad model proves the bad model ran faithfully; it says nothing about whether the output was wise, safe, or lawful. The provenance of the weights, which data trained the model, is an upstream problem, and zero knowledge proofs of training are an earlier, harder research frontier than proofs of inference. And any record still needs a signer: a proof binds a computation to a fingerprint, while accountability binds the fingerprint to a legal entity.

Which raises the practical question. The proofs, quotes, and fingerprints exist as small strings. Where do they live so that an auditor can find them years later, with a timestamp nobody disputes, and no possibility that the operator quietly revised the collection? That is not a cryptography question. It is an infrastructure question, and it is the subject of the final article in this series.

‍

This article is for informational purposes only and does not constitute investment advice.

‍

Mintlayer Web Services provides Bitcoin-native settlement infrastructure for commitments that need to outlive the systems that created them. Learn more →

Discover more

Mintlayer Development Update - September
Development

Mintlayer Development Update - September

This month, development focused on strengthening the security and reliability of Mintlayer tools, hardening the infrastructure that supports them, and preparing the foundation for upcoming products.

September 30, 2026
Mintlayer $ML Migration Update: Final Deadline Confirmed, New Bridge and ERC20 Coming Next
Development

Mintlayer $ML Migration Update: Final Deadline Confirmed, New Bridge and ERC20 Coming Next

The final deadline for migrating the original ERC20 $ML token is confirmed for 1 November 2026 and will not be extended. In parallel, a new permanent bridge and a new ERC20 representation of $ML are on the way.

September 21, 2026
Your Address Checks Itself
Research

Your Address Checks Itself

One typo in a bech32 address gets caught, located, and corrected before signing. On 0x chains, almost any lowercase string is a valid address. Security at Mintlayer starts at the format level.

September 14, 2026
Explore all