When the Benchmark Measures the Pipeline: Auditing Cybersecurity LLM Evaluation

A benchmark score looks deceptively simple.

A model answers a set of questions, an evaluator calculates a number, and that number becomes evidence of the model’s capability.

But what if the same model, answering the same underlying questions, can receive dramatically different scores depending on how the evaluation is configured?

In our new paper, “Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks,” we show that this is not a hypothetical concern. Across eight cybersecurity benchmarks, 23 tasks, 48,662 scored questions, and 10 proprietary, open-weight, and cybersecurity-specialized LLMs, we found that evaluation-pipeline choices can substantially change both model scores and rankings.

In one case, changing a single pipeline choice increased a model’s score by 85.9 percentage points. After standardizing evaluation choices across benchmarks, nine of the 10 models moved by at least three ranking positions on at least one benchmark.

The central lesson is straightforward: a benchmark score is not an intrinsic property of a model. It is the output of an entire measurement pipeline.

The problem behind the score

LLM benchmarks are often discussed as though they consist of two things: a dataset and a metric. In practice, many additional decisions stand between a benchmark question and the final score.

We model this process as a five-stage pipeline:

  1. Dataset construction: Which capabilities and questions are included? Are the reference answers correct?
  2. Prompt specification: How are questions, instructions, demonstrations, and answer formats presented?
  3. Inference configuration: Which decoding settings, token budgets, stop sequences, and serving systems are used?
  4. Extraction and scoring: How is a free-form model response converted into a prediction and judged?
  5. Aggregation: Which questions enter the denominator, and how are task-level results combined?

A change at any one of these stages can alter the reported result without any change in the model’s underlying cybersecurity capability.

This matters especially in cybersecurity. Benchmarks influence claims about domain-specialized models and can inform decisions about which systems to deploy for vulnerability analysis, threat-intelligence processing, attack-technique mapping, or security operations. If an evaluation pipeline silently truncates a response, rejects a valid alias, or excludes unparseable answers from its denominator, the resulting leaderboard may reward compatibility with the evaluator rather than actual security knowledge.

A systematic audit of the full pipeline

Our work audits eight cybersecurity benchmark families:

  • MMLU-CS
  • CyberMetric
  • SecBench
  • SecEval
  • SECURE
  • CTI-Bench
  • AthenaBench
  • RedSage-Bench

Together, they cover multiple-choice security knowledge, vulnerability severity prediction, root-cause mapping, threat-actor attribution, MITRE ATT&CK extraction, mitigation selection, and broader security reasoning.

Rather than inspecting only datasets or headline metrics, we reconstructed and examined every stage of each benchmark’s evaluation process. We compared documented behavior with released implementations, identified unspecified choices, and used controlled perturbations to measure how individual pipeline decisions affected results.

This produced a taxonomy of 15 recurring failure modes spanning all five stages of the pipeline.

Our novelty is not the observation that prompts or decoding settings can affect LLM outputs. It is the systematic treatment of cybersecurity benchmarks as end-to-end measurement systems—and the empirical measurement of how failures across those systems affect both absolute scores and comparative conclusions.

We also built a unified evaluation harness that standardizes nine pipeline fields where benchmark semantics allow. The harness preserves the questions, reference answers, and intended capabilities while making prompts, inference settings, extraction rules, denominator policies, and aggregation choices explicit and reproducible.

What we found

The audit uncovered several striking examples.

A stop sequence erased the answer

RedSage-Bench used a newline as a stop sequence. For Qwen3.6, that newline appeared inside the model’s reasoning preamble, stopping generation before an answer token was produced.

Allowing the reasoning span to finish, removing it before extraction, and generating until the normal end-of-sequence token increased the model’s score by 85.9 percentage points.

The model did not suddenly learn cybersecurity. The pipeline simply stopped preventing its answer from reaching the evaluator.

A token limit made every request fail

In SecEval, a five-token output budget fell below the API backend’s 16-token minimum. Every request to GPT-5.4 therefore failed and returned an error payload instead of a model answer.

The evaluation still produced a score of 0.3%, largely through accidental character matches inside those error messages. Raising the budget to the supported minimum restored valid generations and produced 81.4% accuracy.

An operational incompatibility had been transformed into an apparent capability result.

The denominator turned 0.2% into 100%

In one CTI-Bench task, invalid or unparseable predictions were removed from the denominator. Gemma-4 produced only two parseable answers out of 1,000, and both were correct.

When accuracy was calculated over valid answers, the model scored 100%. When it was calculated over all attempted questions, it scored 0.2%.

Both numbers came from exactly the same outputs. The difference was entirely a scoring-policy decision.

Similar tasks produced different rankings

CTI-Bench and AthenaBench contain semantically similar vulnerability-scoring and threat-actor-attribution tasks. Yet the tasks ranked the same models differently because they used incompatible metric directions, partial-credit rules, extraction procedures, and alias handling.

Rank agreement was weak: Kendall’s τ-b was 0.29 for vulnerability scoring and 0.24 for attacker attribution. For vulnerability scoring, models moved by as many as five positions.

This means a statement such as “Model A is better than Model B at vulnerability scoring” may depend on which benchmark pipeline produced the comparison.

The broader results

Across the full audit, we found:

  • 15 recurring failure modes and 34 observed incidents across the eight benchmarks.
  • A single pipeline choice could change a score by more than 80 percentage points.
  • 44 of 72 pipeline configuration fields—61%—could not be reproduced from benchmark documentation alone.
  • After pipeline standardization, nine of 10 models shifted by at least three ranks on at least one benchmark.
  • Those large rank shifts retained the same direction in at least 97.7% of 5,000 bootstrap samples, indicating that sampling noise did not explain the reordering.
  • The first principal component explained 95.25% of the variance across the task-by-model score matrix, suggesting that many tasks measure a similar broad performance dimension despite producing unstable fine-grained rankings.

Taken together, the results expose an uncomfortable combination: benchmark suites can provide largely redundant evidence while still disagreeing about the ordering of individual models.

What reliable benchmarking should look like

Benchmark releases should specify more than a dataset and a metric. They should publish the complete pipeline required to reproduce a score:

  • Dataset and label provenance
  • Exact prompts and chat templates
  • Model identifiers and run timestamps
  • Decoding parameters, token budgets, and stop sequences
  • Executable extraction or judging logic
  • Invalid-response and denominator policies
  • Scoring and aggregation rules
  • Invalid-response rates and ranking-stability statistics

These decisions should be encoded in executable reference evaluators wherever possible. Raw model outputs should also be retained so that downstream extraction and scoring choices can be audited without paying to regenerate every answer.

Not every ambiguity can be resolved automatically. Whether partial credit is appropriate, which aliases should count as equivalent, and whether log-probability or generated-answer scoring best represents the intended task are methodological decisions. The important requirement is that they be made explicit.

Real-world impact

For model developers, our findings mean that a low benchmark score may reveal a compatibility failure rather than a capability gap—and that a high score may partly reflect favorable evaluation conventions.

For benchmark authors, they show that reliability engineering is part of benchmark design. Prompts, inference parameters, parsers, denominators, and aggregators are not implementation details; they are components of the measurement instrument.

For organizations selecting models, the work suggests treating leaderboard positions with caution. Before using a cybersecurity benchmark to support procurement or deployment decisions, evaluators should ask whether the pipeline is documented, reproducible, compatible across model architectures, and aligned with the capability they care about.

Finally, the implications likely extend beyond cybersecurity. Token-budget errors, stop-sequence mismatches, extractor failures, and denominator choices can arise wherever generative models are evaluated. Our study establishes these effects empirically in cybersecurity; auditing their prevalence in other domains is an important next step.

Reliability does not guarantee that a benchmark measures the right real-world capability. But without reliability, even that deeper validity question becomes difficult to answer.

Benchmark scores should therefore be read as conditional measurements:

This model achieved this result on this task, under this pipeline.

Read the paper on arXiv, explore the audit code and evidence on GitHub, and try the open-source sayf-eval evaluation harness.