Knowing the Rules, Applying the Rules: Evaluating Language Models on Traditional Chinese Bazi

BaZi2500: Task-specific evaluation of rule knowledge and contextual interpretation in Chinese Bazi QA.

arXiv:2610.05682 · 2026

Final-set results

Recovered aggregate results, checked against the manuscript. Complete per-item responses are not included. Scores and intervals describe an outcome-informed final set.

Four-choice baseline 25.0% Best 75.6% vs baseline +50.6pt Questions 2492

Open leaderboard →

Evaluation Protocol

  • Denominator: all final-set records for each endpoint; each native run contains 2,492 items (1,454 Theory and 1,038 Case).
  • Invalid responses: empty, malformed, or multiple-answer final outputs count as wrong and remain in the denominator.
  • Primary aggregation: item-weighted micro accuracy, not equal-weight averaging of Theory and Case. Task columns use their own denominators.
  • Intervals: recovered nominal 95% Wilson score intervals, with z = 1.96. Outcome-informed selection and source/chart clustering are not accounted for; the intervals are descriptive.
  • Random baseline: 25% for four-choice questions. Per-row grey annotations show the accuracy difference from this baseline in percentage points.
  • Data scope: recovered aggregate CSV snapshots. Category-level counts reproduce the headline scores, but complete per-item response records are not included. Configuration contrasts in the figures use the original 3,000 items, not the selected final set.
  • Source: manuscript Evaluation Setup, main results, and Appendix Statistical Definitions; original CSV snapshots are preserved under data/sources/.