Vitalik Buterin tests three-layer privacy setup for querying frontier AI
In brief
- Vitalik Buterin described a personal self-experiment in an X post, U.Today reported.
- Qwen 3.8 Flash Next, a local model, wrote queries and called remote frontier models as tools.
- zkAPI shielded the payment channel; Tor handled network and IP-level privacy.
- The setup worked, Buterin said, but Tor latency was 10-100x higher than it could be.
- Local model speed of 20-30 tokens per second was too slow for comfort, he said.
This was a personal self-experiment, not a product launch.
Three layers of privacy
Buterin used a local model, Qwen 3.8 Flash Next, to orchestrate the process. It called powerful remote models as tools whenever it needed higher-level thinking or knowledge it didn't have itself.
The privacy came in three layers, according to his post. In the first, the local model (not Buterin) wrote the queries sent to the frontier model, so that neither personally identifiable information nor his writing style would give away who he was. A skill file taught it when and how to build requests that revealed minimal data. The second layer used zkAPI to keep his identity from being exposed through the payment channel. The third used Tor for network and IP-level privacy, and Buterin said he ran zkAPI wrapped with Tor as a command-line tool.
"You need all three (and finally we have all three, at least to some extent)," Buterin wrote.
What worked, and what didn't
It worked, he said. The recommendations came back, and Buterin said information from the frontier models helped improve them.
He didn't stop there. Buterin said Tor isn't optimized for request-by-request de-linking, which he called the only form of network-layer privacy that makes sense today, and he judged the Tor setup probably not private enough while carrying latency 10-100x higher than it could be. He also said the skill file's request construction strategies were far from optimal.
Speed was the other sore spot. Qwen 3.8 Flash Next ran at 20-30 tokens per second, which Buterin considered too slow; he said it'd only start to feel fast at 100+.
The core tradeoff
Buterin identified a tradeoff that sits under the whole design. The more carefully a user limits the data handed to a remote model, the less that model can help.
The setup, its results and its weak points are all Buterin's own account, as he described them in his X post and as U.Today reported them.
Frequently asked questions
How did Vitalik Buterin keep his data private from frontier AI models?
Buterin said he used three layers. A local model wrote the queries so his personal details and writing style stayed hidden. zkAPI kept his identity out of the payment channel, and Tor provided network and IP-level privacy.
What problems did Buterin report with his private AI setup?
Buterin said Tor isn't optimized for request-by-request de-linking and that the setup probably wasn't private enough, with latency 10-100x higher than it could be. He also called the skill file's request strategies far from optimal and said the local model's 20-30 tokens per second was too slow.
What tradeoff did Buterin identify in privacy-preserving AI?
Buterin said that the more carefully a user limits the data given to a remote model, the less that model can help. In his setup, a skill file taught the local model to build requests that revealed minimal data.


