Knowing the Rules, Applying the Rules: Evaluating Language Models on Traditional Chinese Bazi
BaZi2500: Task-specific evaluation of rule knowledge and contextual interpretation in Chinese Bazi QA.
Abstract
Knowing domain rules does not guarantee applying them to a case. We study this distinction in traditional Chinese Bazi through 3,000 Chinese multiple-choice questions spanning 14 Theory and 11 Case categories. Six endpoint systems are evaluated, with primary results reported on a 2,492-item model-informed refinement. Theory accuracy exceeds Case accuracy for every system, and gaps of 16.60–29.56 percentage points remain when invalid responses are excluded. The contrast is more specific than a general case-reasoning deficit. Across six systems, Twelve Stages and Nayin reach mean accuracies of 89.10% and 88.62%, while Shensha Basics reaches 75.96%. Within Case, Luck Pillars averages 84.62%, but Career and Family Relations average only 36.98% and 38.19%. Overall rankings also conceal different category strengths. On the original 3,000 items, paired DeepSeek native/disabled comparisons associate native configurations with Theory gains of 6.53 and 12.20 points for Flash and Pro, respectively; Case changes are -3.67 and +1.27 points. These are provider-configuration associations, not isolated causal effects of reasoning. The results motivate task-specific evaluation of cultural-domain applications rather than reliance on aggregate knowledge scores. The benchmark measures agreement with a model-generated, model-verified answer key, not real-world predictive validity. Final-set results are post-selection descriptions, and incomplete provenance and expert validation constrain their interpretation.
Final-set results
Recovered aggregate results, checked against the manuscript. Complete per-item responses are not included. Scores and intervals describe an outcome-informed final set.
Evaluation Protocol
- Denominator: all final-set records for each endpoint; each native run contains 2,492 items (1,454 Theory and 1,038 Case).
- Invalid responses: empty, malformed, or multiple-answer final outputs count as wrong and remain in the denominator.
- Primary aggregation: item-weighted micro accuracy, not equal-weight averaging of Theory and Case. Task columns use their own denominators.
- Intervals: recovered nominal 95% Wilson score intervals, with z = 1.96. Outcome-informed selection and source/chart clustering are not accounted for; the intervals are descriptive.
- Random baseline: 25% for four-choice questions. Per-row grey annotations show the accuracy difference from this baseline in percentage points.
- Data scope: recovered aggregate CSV snapshots. Category-level counts reproduce the headline scores, but complete per-item response records are not included. Configuration contrasts in the figures use the original 3,000 items, not the selected final set.
- Source: manuscript Evaluation Setup, main results, and Appendix Statistical Definitions; original CSV snapshots are preserved under data/sources/.
Key findings
Knowledge and application differ
Every system scores higher on Theory than Case.
Category profiles matter
The useful application-facing conclusion is the profile of limits, not the name of a winner.
Configuration gains depend on the task
Both Theory contrasts are positive and detected, yet neither system shows a correspondingly detected positive Case difference.
Results in detail
Open full-size figure ↗
Open full-size figure ↗
Open full-size figure ↗Limitations
Representation and validation. The model-generated, model-verified key lacks comprehensive expert adjudication. Contested school conventions, incomplete public item provenance, concentrated Case sources, and unknown contamination limit interpretation. Scores measure agreement with this representation, not metaphysical truth, predictive validity, or general cultural competence. Applications involving consequential real-world decisions are outside the supported scope.
Selection and statistical scope. Final items were selected using the evaluated panel, not held out independently. Categories and Case tags are observational groupings; source and chart reuse remain unmodeled. Point estimates and nominal intervals do not establish population-level category effects. Shared generation/evaluation model families and single runs at temperature zero add uncertainty.
Endpoint limitations. Mixed routes, missing immutable snapshots, and configuration offsets restrict capability and efficiency comparisons.