The Honest Quorum Problem: Epistemic Byzantine Fault Tolerance for Agentic Infrastructure
State machine replication and Byzantine fault-tolerant consensus guarantee agreement despite a bounded number of arbitrary faulty participants; Byzantine participants may coordinate or collude. These guarantees rely on participants outside that set correctly implementing the protocol’s transition or validation semantics. Agentic validators expose a weaker partition: an authenticated, responsive, correctly signed, nonequivocating reasoning participant that is protocol-compliant with the voting protocol may nevertheless endorse a semantically invalid state transition. We call the resulting failure mode an epistemic fault and the collective phenomenon the Honest Quorum Problem; here, honest means protocol-compliant, not semantically correct. Such a quorum can satisfy ordinary protocol checks while forming a well-formed certificate for an invalid transition. We show that agreement alone does not establish semantic certificate validity or execution safety. Agentic validators may share model weights or lineage, training distributions, prompts, retrieval sources, toolchains, evidence, reasoning scaffolds, and provider infrastructure, yielding correlated epistemic faults.
Introduction. An operator proposes an infrastructure state mutation: expand a deployment service account from read-only inspection to write authority over a production namespace. Several reasoning validators inspect the same canonical request, state snapshot, policy context, and evidence package. Each validator authenticates correctly, receives the same request digest, follows the voting protocol, signs the expected message, responds before the timeout, and does not equivocate. A quorum approves the transition, but the transition violates an application invariant: it crosses the intended control-plane isolation boundary. The protocol succeeds in forming agreement, yet the system commits a semantically invalid action. Classical state machine replication (SMR) and Byzantine fault-tolerant (BFT) consensus already permit arbitrary faulty replicas to coordinate, collude, equivocate, and choose adversarial messages [1–3].
Discussion / Conclusion. The threshold theorems in Section 6 mix deterministic protocol assumptions with statistical semantic assumptions. Agreement is the deterministic part: under authenticated channels, partial synchrony, and the assumption that only Byzantine validators equivocate, two conflicting q-certificates cannot both form when the intersection condition holds. Semantic certificate validity and liveness are different. They are conditioned on the events Eδ and Uε: the former bounds protocolcompliant false endorsement of invalid candidates by eδ, and the latter bounds unusable support for a valid candidate by uε. The scope of these events must be stated with the certificate. A pointwise, per-candidate claim applies to one candidate, state, policy, and evidence package. A domain-conditional claim applies only inside a calibrated workload domain D. An average-case claim averages over a task distribution on that domain. A uniform claim must bound every admissible task in the stated class. These are not interchangeable guarantees. In particular, the paper’s statistical claims do not turn semantic correctness into a deterministic property of consensus.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Do reasoning benchmarks predict model performance in long-horizon workflows?- How are task bindings validated and what does validation cost per task?
- Who validates task bindings and how is validation checked?
- Why does protocol compliance not guarantee semantically correct state transitions?
- Does semantic validity across a quorum require new property definitions?
- How can we detect when protocol-compliant validators certify semantically incorrect states?
- Why does validator consensus solve agreement but not answer correctness?
- How do false endorsements and unusable support bound validator consensus properties?
- How much do shared prompts and evidence channels correlate validator outputs?
- Can raising the quorum threshold alone fix the correlated faults problem?
- Which diversity targets matter most to reduce validator correlation?
- What evidence would prove validators are independent versus sharing a cause?
- Which shared channels cause the strongest correlated validator failures?
- Can validators gather evidence independently without raising disagreement costs?
- Can a quorum of protocol-compliant validators certify semantically invalid transitions?