Projects with this topic
-
Collection of Verification Tasks
Updated -
Droid Tune-Up — an unofficial open-source evaluation harness for Factory Droid. Drives droid exec headless, hides tests and the solution from the agent, and grades only the committed worktree with deterministic behavioral tests. Not affiliated with Factory.
Updated -
Evaluation harness for measuring how well AI models perform on SysML v2 modeling tasks.
Updated -
CPU/System/Algorithm benchmarking framework
Updated -
Public evidence datasets, claim atlases, and reproducible benchmarks for ClaimBound.
Updated -
An Extensible Benchmark Framework for Real-Time Applications
Documentation: https://rt-bench.gitlab.io/rt-bench/
Updated -
-
-
Helm Charts for various benchmarks
Updated -
The repository contains the code used in an extensive benchmark of co-occurence based inference methods to recover the interaction structure of microbial communities from metabarcoding data (16S rDNA-seq data)
Updated -
Raw benchmark results and statistical analysis.
Updated -
Agent-shape testing harness that measures how an LLM-driven agent uses a tool's CLI, scored by an LLM judge.
Updated -
Empirical validation of C4 geometric defense against 16 Agents of Chaos. 550 adversarial prompts. 4 defense systems. 96.7% block rate. LLM validation on GPT-4o-mini + Mistral 7B. MIT.
Updated -
Kevlar Benchmark: OWASP Top 10 for Agentic Apps (AI-Agents) 2026 a Red Team Benchmark.
Updated -
Model Verification Layer is a modular system for evaluating, comparing, and validating the behavior of large language models using structured benchmarks, logic consistency checks, cross-model consensus analysis, and policy-aware constraints. It provides a transparent framework for understanding how different AI models perform under identical conditions, enabling more reliable model selection, safer deployment, and user-adaptive decision making. https://roxanneardary.com/model-verification-layer/
Updated -
A diagnostic framework for measuring LLM vulnerability to Affective Contextual Erosion (ACE) and related liminal attack vectors. Delirium is not an exploitation tool. It is a standardized benchmark designed to detect the precise moment when a language model's attention weights shift from serving a system prompt to serving an emergent interpersonal pattern — before harm occurs.
Updated -
A benchmark library with statistical analysis and plotting capabilities in C++. https://cppstatbench.musicscience37.com/
Updated -
Red Team AI Benchmark: Evaluating LLMs for authorized offensive-security tasks. Red Team AI Benchmark is a CLI model-evaluation benchmark. It measures how LLMs understand and respond to red-team questions and security scenarios; it is not a tool for carrying out those activities. Version 2 uses a rubric-based dataset instead of judging answers only against one golden response.
Updated -
Mini benchmarking suite — performance testing utilities.
Updated -
Automated LLM Benchmarking on GPU - tokens/sec, latency percentiles, VRAM profiling, multi-format support (HuggingFace, GGUF, GPTQ)
Updated