|
Description
|
Large Language Models (LLMs) have demonstrated remarkable capabilities in mathematical and scientific reasoning tasks, yet standard benchmark evaluations fail to capture the stability of these reasoning processes under semantically equivalent input variations. We introduce a metamorphic testing framework for assessing semantic invariance in LLM reasoning, systematically evaluating model responses across eight metamorphic relations: identity, paraphrase, fact reordering, expansion, contraction, academic context, business context, and contrastive formulation. We conduct a comparative analysis of two recent foundation models - Hermes-4-70B and DeepSeek-R1-0528 - across 79 reasoning problems spanning eight categories (Physics, Mathematics, Chemistry, Economics, Statistics, Biology, Calculus, and Optimization) at three difficulty levels. Both models exhibit similar aggregate invariance scores (~0.90), yet manifest distinct vulnerability profiles-DeepSeek shows notable sensitivity to paraphrase and fact-reordering transformations, particularly in formal domains (Calculus: -0.104, Statistics: -0.048). Contrastive transformations prove particularly challenging for Hermes, inducing mean score deltas of -0.147 (Hermes) and -0.005 (DeepSeek). Semantic similarity analysis further reveals that Hermes maintains more coherent reasoning traces (mean similarity: 0.87) compared to DeepSeek (0.78), with DeepSeek exhibiting near-complete reasoning breakdown in specific category-transformation combinations. These findings demonstrate that metamorphic testing uncovers robustness characteristics invisible to conventional accuracy metrics, providing essential insights for deploying LLMs in high-stakes reasoning applications. We release our evaluation framework and complete experimental data to facilitate reproducible robustness assessment of foundation models. (2025-12-19)
***This entry has been automatically imported via Infodoc(ASO) CSV by LIST harvest scripts. Please refer to https://doi.org/10.1109/ACCESS.2025.3646270 for the original and latest version of the dataset and data downloads*** (2026-06-09)
|