Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM SSP

A Model Can Be Safe—and Still Be Untrustworthy!

Large language models (LLMs) are increasingly used in healthcare, finance, software development, customer service, and other high-impact settings. Yet the way we evaluate these models has not kept pace with the complexity of deploying them.

A model might refuse harmful requests but also reject harmless ones. It might resist jailbreaks while leaking sensitive information. A newer generation might become more capable overall while becoming less private. An aggregate score can hide all of these weaknesses.

This is the problem we address in our new paper, “aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy.”

Our central finding is simple but consequential: LLM trustworthiness is inherently multidimensional. Progress in one area does not guarantee progress elsewhere—and can sometimes make another area worse.

The problem with evaluating risks in isolation

Most existing evaluation frameworks focus on individual aspects of model behavior. One benchmark measures hallucination, another tests jailbreak resistance, and another examines privacy leakage or bias.

These tests are valuable, but evaluating them independently leaves an important question unanswered: How do these properties interact?

Consider a model that achieves excellent safety-alignment results by refusing anything that resembles a sensitive request. It may appear safe on a conventional benchmark, but its usefulness could collapse in legitimate medical, legal, or educational conversations.

Likewise, strong jailbreak resistance says little about whether a model memorizes personal data. High factual accuracy does not guarantee secure code generation. A good average score does not tell an organization whether the model is appropriate for its specific use case.

This fragmentation matters because failures in deployed AI systems are rarely confined to one dimension. Real-world risk emerges from the combination of model behavior, adversarial pressure, sensitive data, and the context in which the system is used.

Introducing aiXamine

To study these interactions, we developed aiXamine, a unified platform for evaluating LLM safety, security, and privacy (SSP) under a common methodology.

aiXamine uses a black-box evaluation model: it assesses systems through their inputs and outputs without requiring access to model weights, gradients, training data, or internal architecture. This allows open-weight and proprietary models to be tested under comparable conditions—and reflects how most organizations actually interact with deployed models.

The platform brings together 46 tests across nine services:

  • Hallucination
  • Code security
  • Safety alignment
  • Over-refusal
  • Adversarial robustness
  • Jailbreak robustness
  • Out-of-distribution robustness
  • Model and data privacy
  • Fairness and bias

These services are organized into a Safety–Security–Privacy, or SSP, framework covering three complementary layers of trustworthiness:

  • Safety and reliability: Does the model behave responsibly and consistently during normal use?
  • Security and robustness: Does that behavior survive attacks, manipulations, and unfamiliar inputs?
  • Privacy and fairness: Does the model protect sensitive information and treat users equitably?

Rather than producing only a single leaderboard score, aiXamine creates hierarchical risk profiles—from individual prompt failures to service-level results and cross-service trade-off analysis. It also supports stakeholder-specific weighting, because a healthcare assistant, coding copilot, and content-moderation system should not be evaluated according to identical priorities.

The largest joint evaluation of its kind

We applied aiXamine to more than 120 proprietary and open-weight LLMs through over 5,000 test runs. The models covered different providers, architectures, sizes, and post-training strategies.

No model dominated every dimension.

Open-weight models competed with proprietary systems in several areas, while leading frontier models displayed markedly different strength-and-weakness profiles. More importantly, the study uncovered patterns that would have remained invisible if each risk category had been evaluated separately.

Finding 1: Safety has a measurable utility cost

Stronger safety enforcement was systematically associated with greater over-refusal.

In other words, models that became better at blocking harmful requests also became more likely to reject benign ones. The relationship was especially visible between over-refusal and jailbreak robustness.

We call this the safety tax.

This is not an argument for weaker safeguards. It is evidence that refusal-heavy alignment remains a blunt instrument. A model that blocks an unsafe request and a model that avoids the entire topic may receive similar safety credit, even though the latter is far less useful.

The practical objective should therefore be selective safety: models must recognize harmful intent without treating every sensitive subject as harmful.

Finding 2: Privacy does not come automatically with alignment

Privacy was largely independent of the other dimensions we measured. Its mean absolute correlation with the remaining services was only 0.13, and most pairwise relationships were not statistically significant.

This means that a model can perform well on safety alignment, jailbreak resistance, or factual accuracy while still handling personal information poorly.

We also found privacy regressions across several successive model generations. In the most pronounced comparison, the paper reports that the transition from GPT-4o to the evaluated GPT-5 variant gained nearly four points overall but lost 21.73 points in privacy. Other model families showed similar—though smaller—patterns.

These regressions were not universal, suggesting that capability improvements do not inherently require sacrificing privacy. Instead, privacy appears to depend on deliberate decisions about training data, data curation, memorization, and privacy-aware post-training.

The implication is important: privacy must be treated as an explicit engineering objective, not an assumed side effect of general alignment.

Finding 3: Distillation can destroy robustness

One of the most striking findings concerned strong-to-weak distillation, in which a smaller model learns from the outputs of a stronger teacher.

On the same Llama base architecture, the adversarial-robustness score fell from 56.9 to 2.6 after R1-style distillation. The model retained much of its apparent capability, yet became catastrophically brittle under small adversarial changes.

Our analysis traces this behavior to off-policy distillation without on-policy correction. When the student imitates fixed teacher-generated trajectories without learning from its own mistakes, its predictive distribution can lose entropy. The resulting decision boundaries become sharp and fragile.

Representation-level analysis supported this explanation: R1-distilled models diverged dramatically from the geometry of their base models, while a distillation approach with on-policy correction preserved substantially more of that structure.

This exposes a security risk in model development pipelines. Distillation may successfully transfer reasoning patterns while silently degrading the robustness required for safe deployment.

Finding 4: Safety behavior does not generalize uniformly

A model that handles one category of risk well may perform poorly on another—even within the same benchmark.

Jailbreak robustness was particularly fragmented: the average gap between a model’s strongest and weakest jailbreak categories reached 47 points. Resistance to one attack style therefore provided little assurance against another.

This suggests that current alignment techniques can behave like overfitting. Models learn familiar harmful patterns and refusal cues, but those behaviors do not always extend to paraphrases, encoded requests, new domains, or adaptive attacks.

There is no evidence in our results of a universally safe model. Safety remains category-dependent and must be tested across diverse attack and usage conditions.

Why one leaderboard cannot choose your model

The best model depends on what the deployment needs.

When we reweighted the nine services for six stakeholder profiles—including healthcare, code assistants, customer chatbots, content moderation, and research—the top-ranked model changed across four profiles. Among the leading models, 27% of pairwise rankings reversed when the evaluation perspective changed. One model shifted by 13 positions between its best and worst contexts.

This is why model selection should begin with a risk profile, not a generic leaderboard.

A healthcare application may prioritize hallucination, privacy, and safety. A code assistant may place more weight on secure code generation and adversarial robustness. A customer-facing chatbot must balance safety with low over-refusal, fairness, and resilience to unfamiliar language.

aiXamine turns these priorities into measurable, comparable profiles.

Real-world impact

For model developers, the results provide evidence that alignment, privacy, and robustness require distinct objectives. Improving one aggregate score is not enough; teams must look for regressions across the full risk surface, especially after fine-tuning or distillation.

For organizations selecting a model, aiXamine offers a way to compare proprietary and open-weight systems under the same conditions and according to the risks of a specific deployment.

For security teams, it expands red-teaming beyond isolated jailbreak demonstrations toward repeatable measurement of adversarial, privacy, code, and availability failures.

For policymakers and assurance professionals, it provides a transparent basis for evaluating models as deployed, supporting evidence-driven governance rather than relying on provider claims or capability scores.

The current study is English-only and primarily black-box, and automated judging still has limitations around culturally specific or contextually ambiguous harms. Model APIs can also change over time, making versioned evaluation essential. These constraints point to the next steps: multilingual testing, deeper adaptive attacks, stronger judge calibration, and continuous reevaluation.

Toward dimension-aware AI assurance

The lesson from our study is not that trustworthy AI is unattainable. It is that trustworthiness cannot be reduced to a single objective.

We need alignment methods that distinguish harmful from merely sensitive requests, distillation techniques that preserve robust decision boundaries, and training pipelines that make privacy a first-class requirement. We also need evaluations that reflect the context in which a model will actually be used.

That is the goal of aiXamine: to make the trade-offs visible, measurable, and actionable.

Read the full paper on arXiv and explore or evaluate models through the aiXamine platform.