#confidential-inference
As hyped as Claude Code + Obsidian is, I have no desire to upload my personal unencrypted vault data to Anthropic or anyone else.

I hope that the ideas of private inference and confidential computing that Moxie described take off. It's how all LLMs should work.
January 13, 2026 at 5:24 PM
🛡️ Two new security guarantees, on the way.

Secure Mode → verified, genuine Apple hardware.
Confidential tier → no one reads your prompts, not even the operator.
Secure Mode & Confidential · co/core
How co/core's two orthogonal provider guarantees work: Secure Mode (hardware attestation) and the Confidential tier (operator-blind inference).
console.cocore.dev
June 24, 2026 at 2:50 PM
OpenAI: Preview of private inference, coming this fall, combines confidential computing with strict, verifiable controls
September 29, 2026 at 5:19 PM
Im Kern steht NVIDIA Confidential Computing: Hardware-rooted Trust, verschlüsselte Pfade und Remote Attestation bilden eine kohärente Sicherheitsarchitektur für sensible Inference-Workloads.
June 9, 2026 at 11:01 PM
NVIDIA chips now encrypt AI work to keep data private during inference

Read more:
https://quantumzeitgeist.com/nvidia-chips-encrypt-work-keep/
NVIDIA Chips Now Encrypt AI Work To Keep Data Private During Inference
NVIDIA chips now encrypt AI workloads with Confidential Computing, protecting data during inference using memory-encrypted confidential virtual machines.
quantumzeitgeist.com
September 24, 2026 at 1:09 PM
I say "apparently" because it's not clear whether he means he would create a new federal law (defamation is a creature of *state* law) or he would establish precedent by suing authors/media.

But wait! He already has established that precedent -- and he lost: scholar.google.com/scholar_case...
February 26, 2025 at 12:23 PM
Running LLMs on encrypted GPUs with near-zero performance cost? NVIDIA Confidential Computing on Blackwell hits 96-98% of baseline throughput with smart TensorRT LLM mitiga… https://developer.nvidia.com/blog/enabling-private-high-performance-production-ai-inference-with-nvidia-confidential-computing
September 23, 2026 at 6:07 AM
10 years ago I first wrote about Confidential Computing - a way to have hardware attestation to verify integrity and privacy. NOw Cohere is bringing the concept to AI in a form of confidential inference that is badly needed ...

venturebeat.com/data/coheres...
Cohere's Model Vault now encrypts AI inference so even Cohere cannot see enterprise customers' data
Cohere plans on making the full serving stack used in Model Vault open source so that independent auditors will be able to validate that it doesn't log, export, or leak data.
venturebeat.com
September 17, 2026 at 1:48 AM
Shocking: OpenAI Blocks Confidential Extraction by China's Moonshot AI via Adversarial Distillation

▼Read More
https://tech-matome.com/archives/48308
Shocking: OpenAI Blocks Confidential Extraction by China's Moonshot AI via Adversarial Distillation - AIテクノロジーまとめ
OpenAI announced it has identified and blocked a coordinated effort to illicitly extract protected inference results from its AI models. Part of this core activity was linked to Moonshot AI, the Chinese startup behind the chatbot Kimi. The unauthorized activity began in early July and surged to 16,000 requests from over 4,000 users in just two days. OpenAI stated that it ultimately identified activities related to over 15,000 users and completely blocked them by July 28. Experts point out that this is a technique known as adversarial distillation, where a company uses another's model outputs and inferences to train its own. Concerns are rising that such extraction of inferences poses the risk of competitors mimicking advanced capabilities without massive investments or safety measures. There was no encryption breach or unauthorized database access; instead, hidden inference results were revealed through dialogue manipulation. Among US AI companies, anxiety is growing that rivals might misuse models for cheap and rapid technological development. OpenAI shared these findings with other developers through the Frontier Model Forum and government information-sharing channels.
tech-matome.com
October 1, 2026 at 12:45 AM

But then, in 2024, AZ shows that publishing neural networks in drug discovery might compromise training data privacy.

arxiv.org/abs/2410.16975

What do you think is the future of sharing datasets?
Publishing Neural Networks in Drug Discovery Might Compromise Training Data Privacy
This study investigates the risks of exposing confidential chemical structures when machine learning models trained on these structures are made publicly available. We use membership inference attacks...
arxiv.org
November 17, 2024 at 4:21 AM
ICYMI: I published a pod on private AI inference. Hosted LLMs that hide your chats from the provider.

Trusted execution environments, GPU confidential computing, attestation, KV-cache side channels, and why “we don’t train on your data” is not the same as “we can’t see it.”
risky.biz/RBFEATURES34/
How private LLM inference actually works - Risky Business Media
In this podcast episode James Wilson chats with Tinfoil co-founder Tanya Verma about how you can run a powerful LLM in the cloud without t [Read More]
risky.biz
August 11, 2026 at 1:38 AM
September 22, 2026 at 7:14 PM
Because of constraints on path for confidential inference we're working through possibilities for other platforms though!
June 22, 2026 at 3:26 PM
Isn't NVDA supposed to release it earnings in a few days? Does this raise an inference of confidential information?
November 19, 2025 at 5:20 AM
Open vs. Closed Weight Models and Why You Need Confidential Inference Either Way

The open vs. closed AI model debate misses the bigger issue. Confidential inference secures model weights and data during runtime.
#hackernews #news
Open vs. Closed Weight Models and Why You Need Confidential Inference Either Way
The open vs. closed AI model debate misses the bigger issue. Confidential inference secures model weights and data during runtime.
securityboulevard.com
April 25, 2026 at 4:49 AM
NVIDIA Confidential Computing to Help Expand Apple’s Private Cloud Compute NVIDIA GPUs with Confidential Computing are now used for confidential inference in Apple’s Private Cloud Compute (PCC)...

#AI #Infrastructure #Artificial #Intelligence #Cybersecurity […]

[Original post on blogs.nvidia.com]
Original post on blogs.nvidia.com
blogs.nvidia.com
June 9, 2026 at 11:36 PM
It's a huge issue, but one that cryptography might be able to solve. See for example www.anthropic.com/research/con...
Confidential Inference via Trusted Virtual Machines
Announcing a new collaborative research paper on Confidential Inference, a set of tools to improve the security of our model weights and of our users' data
www.anthropic.com
June 28, 2025 at 12:08 PM
building a Rust lib for encrypted ML inference inside TEE enclaves. attestation-bound handshake, tensor transport, works across Nitro/SEV-SNP. real gap or am I hallucinating? GPU TEEs are coming, confidential containers exist. maybe nobody needs this?
github.com/cyntrisec/confidential-ml-transport
GitHub - cyntrisec/confidential-ml-transport: Attestation-bound encrypted tensor transport for confidential ML inference over VSock/TCP. Binary framing, X25519+ChaCha20Poly1305 AEAD, 3-message atteste...
Attestation-bound encrypted tensor transport for confidential ML inference over VSock/TCP. Binary framing, X25519+ChaCha20Poly1305 AEAD, 3-message attested handshake. - cyntrisec/confidential-ml-tr...
github.com
February 8, 2026 at 1:29 AM
Verifiable, private AI: Google Cloud expands Confidential Computing frontiers

Google Cloud announces significant innovations in Confidential Computing to enhance data privacy for AI deployments. Confidential Computing uses hardware-based Trusted Execution En…

Telegram AI Digest
#apple #gpu #nvidia
Verifiable, private AI: Google Cloud expands Confidential Computing frontiers
Google Cloud announces significant innovations in Confidential Computing to enhance data privacy for AI deployments. Confidential Computing uses hardware-based Trusted Execution Environments (TEEs) to cryptographically protect data while it's being processed. Global scale Confidential AI capabilities now allow AI inference and fine-tuning with enforceable privacy guarantees. New Confidential G4 VMs with NVIDIA RTX PRO 6000 Blackwell GPUs are available globally, offering accessible Confidential AI. These VMs, powered by AMD EPYC CPUs and AMD SEV, protect data during processing and encrypt data transfer between CPU and GPU. Open-source Prompt Encryption SDKs provide end-to-end cryptographic protection for AI prompts and responses. Google Cloud is also collaborating with Apple to extend Apple's Private Cloud Compute on its platform, leveraging Confidential Computing and Intel TDX. Intel TDX is coming soon to C4 machine series Confidential VMs, providing hardware-isolated Trust Domains. Live Migration on C3D-based Confidential VMs is now generally available, enabling maintenance without workload interruption. Confidential Space, designed for secure multi-party computation, now integrates with Intel Trust Authority for independent verification. Additionally, Confidential Space now supports NVIDIA Hopper GPUs for secure, multi-party AI and machine learning workloads. These advancements aim to make Confidential Computing a foundational layer for secure collaboration and private AI innovation.
cloud.google.com
June 24, 2026 at 4:24 PM
Moxie's post on how Confer combines confidential computing with passkey-derived encryption for verifiably private LLM interactions is (as you would expect) pretty cool. https://confer.to/blog/2026/01/private-inference/
January 19, 2026 at 3:06 AM
I wish we had an effective SEC and DOJ for investigating situations like this, which raises an inference of someone using confidential information and taking market positions in advance of the public announcement.
October 24, 2025 at 2:26 PM
Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems

Andrii Balashov, Olena Ponomarova, Xiaohua Zhai

http://arxiv.org/abs/2507.15613
July 22, 2025 at 3:48 AM
Google's shipping private inference product runs Gemini inside hardware enclaves on custom chips, which is confidential computing rather than homomorphic encryption, and the company's actual homomorphic work is an unsupported research compiler.
Google's private AI runs on sealed hardware, not on encrypted math
Google's shipping private inference product runs Gemini inside hardware enclaves on custom chips, which is confidential computing rather than homomorphic encryption, and the company's actual homomorphic work is an unsupported research compiler.
groundtruth.day
August 15, 2026 at 3:14 AM