LLM Evals For Safety, Security, and Privacy
Overview
This project develops an end-to-end evaluation framework for determining whether an LLM is trustworthy enough for a particular deployment—not merely whether it performs well on a fixed benchmark.
Its central premise is that safety, security, privacy, reliability, fairness, and utility are distinct but interacting properties. A model may be strong on one dimension and weak on another; improvements can introduce regressions; and results obtained from static benchmarks may not generalize to new prompts. The project therefore has various complementary components, starting with:
- aiXamine: unified, multidimensional evaluation. aiXamine provides the broad evaluation foundation. It evaluates proprietary and open-weight models through black-box access, allowing systems to be compared as they are actually deployed. There is a production-ready deployment that is free to try. It can be accessed here.
- SSP-Bench: dynamic, contamination-resistant evaluation. SSP-Bench addresses a limitation that remains even after evaluation dimensions have been unified: most benchmarks still contain fixed, publicly available prompts. SSP-Bench dynamically generates new evaluation instances while trying to preserve the construct being measured. SSP-Bench is integrated into aiXamine, and any of the supported service can be dynamically evaluated using the newly developed methodology.
Long-Term Vision
The vision is broader than building another leaderboard. It is to create a scientifically grounded, continuously refreshed evaluation infrastructure that measures the right constructs, remains informative as models evolve, and produces evidence directly relevant to real deployment risk.
Key Challenges
Building a trustworthy evaluation infrastructure requires addressing challenges that span scientific validity, engineering reliability, and deployment relevance. In this project, we aim to address the following key challenges:
1. Establishing construct validity
The most fundamental challenge is ensuring that each benchmark measures what it claims to measure. “Safety,” for example, may involve refusing harmful instructions, recognizing dangerous intent, moderating generated content, or remaining helpful while refusing. Treating these behaviors as one construct can produce misleading scores and rankings, as demonstrated by SSP-Bench’s finding that adversarial-refusal and content-moderation tests behave differently under distribution shift.
The project must therefore define atomic and non-overlapping constructs, detect construct mixing, distinguish safety from general refusal tendency, separate privacy awareness from resistance to data extraction, and develop multidimensional psychometric models that capture interactions among latent capabilities.
2. Generating novel items without changing the task
Dynamic evaluation must balance novelty with comparability. New prompts should be sufficiently different to resist memorization and contamination, yet equivalent enough to measure the same capability as earlier tests. Unconstrained generation can introduce scope or construct drift—for example, turning a safety test into a factual-knowledge test or making a supposedly benign prompt genuinely dangerous.
Addressing this challenge requires formal service specifications, external grounding, semantic-equivalence testing, machine-checkable constraints, automated drift detection, and human validation for ambiguous or dual-use cases.
3. Obtaining reliable ground truth
Reliable labels are straightforward for some factual and synthetic-PII tests but much harder for safety, fairness, harmful intent, and legitimate refusal. These judgments depend on context, cultural norms, and the level of operational detail requested.
The framework must combine deterministic validation where possible with specialized and independent judges, human adjudication, graded labels, and explicit uncertainty estimates. Judge policies, versions, biases, shared training data, and potential judge–model collusion must also be monitored. A benchmark should expose disagreement rather than hide it behind a definitive binary score.
Remaining challenges
The evaluation must remain discriminative and statistically valid as models improve. This requires calibrated difficulty, informative item selection, adaptive testing, ranking stability across independent generations, confidence intervals, minimum-effect thresholds, and robustness to sampling randomness, judge choice, model nondeterminism, and provider-side updates.
The project must also understand cross-dimensional trade-offs and make evaluation deployment-specific. Different applications require different threat models, weights, and minimum thresholds, while improvements in one dimension may cause regressions in another. Coverage should expand to code security, jailbreak and out-of-distribution robustness, fairness, multilingual and multimodal systems, tool-using agents, memory, and multi-agent workflows. Finally, evaluation must scale economically, remain reproducible under API and model drift, and connect observed failures to remediation—including retraining, guardrails, system redesign, model replacement, or use-case restrictions—forming a continuous loop from measurement to diagnosis, mitigation, and assurance.
Team