news
- New preprint: Tail-Influence Sampling for CVaR Policy Evaluation. TIS allocates a fixed evaluation budget toward the workflow components that matter most for lower-tail risk.
- New preprint: LeanPolish: Verified Supervision for Lean Proof Compression. A neurosymbolic Lean 4 proof-compression method that records every verified edit, exposes the selection effects of search-generated supervision, and shows when learning from it helps beyond symbolic search. Code and dataset.
- VoiceAgentGuard: Contract-Gated Tool Release for Production Voice Agents accepted at EMNLP 2026. VoiceAgentGuard prevents unsafe private-read and state-changing tool calls in voice agents by forcing every proposed action through an auditable contract gate before execution.
- New blog post: What the agent may do — verified authority envelopes for AI agents in GitHub workflows. A Lean-verified checker decides, for sessions of any length, what a coding agent’s session is permitted to do; the model is validated against the action’s own code and measured across public workflows. Code, benchmark, and corpus.
- CovCal at the ICML AI for Math Workshop: a risk-controlled Lean-as-judge framework for natural-language mathematical reasoning, calibrating when Lean proofs can be trusted as partial evidence for answer selection. Code and project page.