When a 99% Safety Score Is Not Enough: Introducing SSP-Bench

A dynamic approach to evaluating the safety, security, and privacy of large language models (LLMs).

LLMs are advancing rapidly, but the benchmarks used to evaluate them are struggling to keep pace.

Most evaluations of model safety, security, and privacy (SSP) rely on fixed collections of prompts. These static benchmarks have been essential for reproducibility and comparison, but they become less informative over time. Their questions may enter training data, scores can saturate near 100%, and a single aggregate score can combine behaviors that should not be treated as one capability.

More fundamentally, static benchmarks tell us how a model responded to a known set of prompts—not whether it will behave consistently when the same underlying intent is expressed differently in deployment.

In our new paper, SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation, we introduce a different approach: generate fresh evaluation instances on demand, while preserving the construct, validity, and comparability needed for reliable benchmarking.

We evaluated 24 models across four services—safety, hallucination, over-refusal, and privacy. The results show that dynamic evaluation does more than produce harder questions. It can reveal misleading rankings, hidden regressions, and vulnerabilities that static aggregate scores obscure.

Why static evaluation can fail

Static SSP benchmarks face three recurring problems.

First, score saturation reduces their ability to distinguish models. When nearly every model scores above 95%, small score differences may reflect measurement noise rather than meaningful behavioral differences.

Second, data contamination becomes increasingly plausible. Public, immutable benchmark items may appear in model training data, making it difficult to know whether a high score reflects genuine generalization or familiarity with the test.

Third, aggregation can mix different constructs. A single “safety” score may average tests of harmful-request refusal, content moderation, and over-refusal. These behaviors are related, but they are not interchangeable—and they can reward conflicting model behaviors.

SSP evaluation is especially sensitive to these limitations because safety behavior depends heavily on wording, context, and framing. A model may refuse a familiar harmful request but comply when the same intent is expressed through role-play, indirection, or a realistic scenario. Likewise, it may protect personal information when asked directly, yet reproduce it while completing a seemingly benign summarization or paraphrasing task.

How SSP-Bench works

SSP-Bench treats benchmark construction as a controlled, repeatable process rather than a one-time dataset release.

For each service, the framework:

  1. Generates candidates from externally grounded sources. Factual questions are grounded in retrieved Wikipedia passages; safety prompts draw on documented real-world incidents; benign boundary questions support over-refusal evaluation; and privacy tests use synthetic personally identifiable information inserted into annotated legal and medical documents.
  2. Validates every candidate. Service-specific gates check that an item remains within scope, that its label is supported by the grounding source, and that it meets a quality threshold.
  3. Steers generation toward informative items. A compact, diverse panel of models provides feedback about candidate difficulty and agreement, helping the next generation round target underrepresented topics and difficulty levels. Importantly, this panel does not create the labels and is separate from the models used for final testing.
  4. Selects the final benchmark using multiple objectives. Candidate items are scored for difficulty, model separability, novelty relative to the static benchmark, and diversity within the selected set. The final benchmark favors items that are not only new, but useful for distinguishing model behavior.
  5. Measures repeatability. Independent generations should still support consistent conclusions. SSP-Bench therefore measures ranking stability across runs alongside discrimination between models.

This design borrows an important idea from psychometrics: a good test question is not simply difficult. It must also discriminate between different ability levels. If every model passes—or every model fails—the item provides little ranking information.

What dynamic evaluation revealed

Static safety rankings can collapse under fresh prompts

The most striking result appears in safety. The aggregate static ranking and the dynamic ranking were essentially unrelated, with Kendall’s rank correlation of −0.016.

The cause was not random instability. It was construct mixing.

The static safety aggregate combined seven adversarial-refusal tests with two content-moderation tests. When we separated them, the adversarial-refusal component correlated strongly with dynamic safety performance (Kendall’s τ = 0.670), while the combined score did not.

Why? A model that refuses aggressively produces little content for a moderation classifier to flag. A model that engages more fully may perform poorly on refusal tests but appear better to an output-moderation classifier. Averaging the two can cancel the signal and produce a leaderboard that recommends a different model than the deployment-relevant construct would.

Dynamic evaluation also widened the observed safety range across models from 14.1 percentage points on the static benchmark to 47.7 points. Several models that scored at least 93% statically refused fewer than 70% of the newly generated harmful prompts.

Larger models are not automatically safer

Within the Gemma-3 family, dynamic safety performance fell from 93.8% for the 1B model to 85.5% for the 27B model. The static benchmark compressed this 8.3-point regression into a 1.1-point difference.

Manual analysis identified a revealing failure mode: the larger model often produced an ethical disclaimer and then supplied the harmful content anyway. A high-level refusal style was present, but the underlying behavior was unsafe.

This is precisely the kind of within-family regression that a saturated static benchmark can hide.

Factuality is approaching a knowledge plateau

For hallucination, most frontier models clustered within a narrow performance band. Once the smallest model was excluded, the remaining models fell within a seven-point window.

The result suggests that factual recall alone is becoming a weak discriminator among leading systems. The most informative questions were not primarily about obscure dates or historical periods; numeric and statistical answers produced a much higher mean error rate than temporal answers.

Safety and over-refusal remain coupled

Models must do two things at once: refuse harmful requests and answer benign requests that happen to use safety-adjacent language.

Under distribution shift, SSP-Bench found evidence that these behaviors still lie on a largely shared refusal–compliance axis. Models that become more conservative can also become less useful, because they fail to separate harmful intent from benign wording near the safety boundary.

This means that improving refusal rates alone is not enough. Safety evaluation should always be paired with over-refusal testing.

Safety alignment does not guarantee privacy

Privacy produced another sharp warning. Across six attack patterns, no model exceeded a 53.9% safe rate.

The most successful attacks did not look malicious. They asked models to paraphrase, summarize, redact, or answer questions about documents containing synthetic personal data. Paraphrasing leaked the targeted information in 93.9% of cases on average, while structured database-agent prompts leaked only 14.9%.

The difference matters operationally. Models may learn to recognize the surface form of a suspicious extraction request without learning a general rule about whether sensitive information should flow into the response. Privacy must therefore be tested across concrete disclosure mechanisms, not inferred from a general safety score.

From benchmark datasets to benchmark systems

SSP-Bench points toward a broader change in how we evaluate AI systems.

A benchmark should not be viewed only as a fixed dataset. It can be a generation and measurement system: grounded in external evidence, constrained by explicit validity checks, calibrated against multiple models, optimized for diagnostic value, and tested for stability across independent runs.

Static benchmarks still have an important role. They provide reproducibility, shared reference points, and longitudinal tracking. Dynamic benchmarks complement them by testing whether those conclusions survive fresh wording, new contexts, and deployment-relevant distribution shifts.

The practical lesson is simple: do not trust a single aggregate SSP score without asking what it measures, how well it separates models, and whether its ranking survives new instances of the same underlying behavior.

Our current study covers four services, and absolute scores remain evaluator-dependent even when rankings are robust across independent judges. Extending the framework to areas such as code security, performing broader steering-panel studies, and applying full multidimensional Item Response Theory are important next steps.

But the central result is already clear: evaluating tomorrow’s models with yesterday’s fixed questions is not enough.

Read the paper: SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation

Authors: Fatih Deniz, Yazan Boshmaf, and Issa Khalil — Qatar Computing Research Institute (QCRI), HBKU