Verifying AI safety between adversaries presents a paradox: you must prove a system is safe without showing how it works. Traditional inspections, like those used for nuclear weapons, often fail here because code and weights are intangible and easily copied. Instead, we should focus on zero-knowledge proofs and hardware-based telemetry.
Zero-knowledge proofs allow a developer to mathematically demonstrate that a model adheres to specific safety constraints—such as refusing to generate biological weapon instructions—without ever revealing the underlying architecture or training data. This protects intellectual property. Meanwhile, secure hardware enclaves can monitor model outputs in real-time, sending encrypted