On Calibration, Consensus Failure, and Ground Truth Ambiguity in Multi-Agent LLM Systems: A Case Study of the Asymmetric Tribunal Architecture
ATA System Research Paper On Calibration, Consensus Failure, and Ground Truth Ambiguity in Multi-Agent LLM Systems: A Case Study of the Asymmetric Tribunal Architecture (ATA) Abstract Multi-agent large language model (LLM) systems, particularly those employing role-specialized ensembles, promise improved reasoning through structured disagreement and synthesis. However, such systems introduce new classes of systemic failure, including false consensus , calibration drift , and ground truth ambiguity . This paper analyzes the Asymmetric Tribunal Architecture (ATA), a four-agent deliberative system with Brier-based reliability calibration, and identifies its critical weaknesses in real-world deployment. We argue that the core limitation lies not in consensus formation, but in the misuse of binary outcome metrics in domains lacking objective ground truth. We propose a lightweight tri-layer outcome framework and delayed calibration mechanism that prese...