Research prototype

Check what AI did, without trusting chips.

Proofs can show which model answered, and would show how chips were used.

Pace the frontier

Compute limits hold only if use is checked.

Earn customer trust

Customers could check which model answered them.

No special hardware

No enclaves or chip-maker keys, just rented GPUs.

How it works

Six steps, one hall

Planned cluster
  1. Register the GPUs

    Every GPU is listed; the list caps compute.

    Rests onA complete GPU list
    More

    The list sets the ceiling in step 6. Nobody inspects the hall, so a complete list is a premise, not a check.

  2. Admit the model

    The checker fingerprints the approved weights. Every proof must match.

    Would proveThe approved model answered
    More

    Proofs are checked against the checker's own fingerprint; that it fingerprinted the real published checkpoint is a premise. The proof covers an integer version of the model; how closely it matches the released model is not measured yet. A verifier flaw found on 29 Sep is fixed (reviewed), and earlier proofs are being re-checked with the fixed verifier; until then they count .

  3. Record every request

    A recorder seals every exchange in order, so none can be removed later.

    ShowsEvery exchange, in order
    More

    Each entry seals the one before it, so an edit breaks every later seal. Today the log and its signed checkpoints stay inside the recorder: outside witnesses are built but not tested yet. Traffic that skips the recorder isn't covered.

  4. Prove random answers

    After sealing, random answers are picked and proven.

    Would proveAnswers weren't edited
    More

    Picking after sealing means the operator can't know which answers get checked. Our zero-knowledge proof of a 7B model took on an 8-core cloud CPU.

    How soon a cheater is caught, 99 times in 100
    Answers provenFake answers it takes
    1 in 100459
    1 in 1,0004,603
    1 in 10,00046,050
    44 proofs catch a batch that is 10% fake, 99 times in 100, however large.
    If answers are sealed before an unpredictable pick, and each proof is sound.
  5. Check where it ran

    A nearby landmark would time each reply. Light is only so fast, so a quick reply means nearby work.

    Would placeWork within
    More

    Light covers per millisecond of round trip, so a deadline allows at most . No landmark exists yet: these are sizing numbers, and the landmark's position would be trusted.

  6. Balance the books and publish

    Proven work is weighed against capacity. The gap would be published.

    BoundsUnexplained capacity
    More

    The bound holds only if its premises hold: a complete GPU list, the recorder as the only way in, and work and capacity counted in the same units. Publishing to outside witnesses is built but not tested yet.

What they don’t prove
  • That the GPU list is complete. The rental record is trusted.
  • That the model is safe, or what it was trained on.
  • How closely the proven integer version of the model matches the released model. That is not measured yet.
  • Anything about traffic that goes around the recorder.
  • Answers that weren't picked are covered only statistically, and today the proofs are being re-checked after a verifier fix.
  • Which GPU did the work. The landmark's position would be trusted, and no landmark exists yet.
  • Compute outside the declared cluster. The bound holds only if its premises hold.
  • That the parties are independent. In this demo Lucid would run every role; today only the recorder and verifier software exist, and Lucid runs both.

Every limit, with its conditions

Try it out

As a user

  • Know which AI answered you
  • See what your agent may do
  • Receipts anyone can check

Checkable claims about AI, without trusted hardware