V1, V2, and Codex Evaluation Results¶
Status: Development retrieval smoke complete; formal comparison not started V1 baseline: Available in the existing Bauer Retrieval Benchmark V2 result: Retrieval improves, but exact lookup misses the provisional promotion gate Promotion decision: Continue private V2 evaluation; retain normal Bauer on V1
These are measured development retrieval results produced under the RAG V2 Evaluation Protocol. They validate the deployed private path but do not substitute for the required verified-gold, five-run, end-to-end, frozen-evidence, and blind-holdout comparison.
Run identity¶
| Item | Value |
|---|---|
| Evaluation date | 2026-07-23 |
| Run | retrieval-smoke-20260723-05; 30 development cases, one V1/V2 repetition |
| Corpus manifest checksum | 2934bb3bf75514a4ae4d340cb68c023831ca20a5e89e194fd20774b1e561b0d8 |
| V1 source baseline | 2815a8f2ba1b7db04d40a80934544f092e54518b |
| V2 evaluated source | bf3c3cb0a96f054ad064824623baf25b4785005b |
| RAG / LibreChat deployments | 4c779152-def4-4c9a-a245-3e205a603547 / 1142dfd1-fc50-477a-8cc4-7f226cc212c2 |
| V1 / V2 Agent IDs | agent_Z8A2LtQWLeP4KuUDbvqZL / agent_pmPMcA25UXS7vznUaz-DU |
| V2 index | bauer-rag-v2-2026-07; run 7e23f6e1-d876-4da0-a677-9a3cdb37db36 |
| Embedding | local/embed-engineering-1024 |
| Reranker | deterministic-multilingual-fallback-v1; remote not configured |
| Answer model | local/qwen-coder; not exercised by this retrieval-only smoke |
| Evaluation artifact commit | d25468e |
| Raw-run / score SHA-256 | fdaf707c0ddf5000712ef176b583b07ae2030fc110e4a0cbe2579177ac534d4a / b929e17c24a7955cd8ec0457898d22a6121e5cf6420aa877df24281e62330a5b |
Promotion gates¶
| Gate | Required | Measured | Result |
|---|---|---|---|
| Exact identifier/document lookup | >=95% | 66.67% | Fail against provisional locations |
| Exact-answer accuracy | >=90% | Pending | Pending |
| Citation correctness | >=95% | Pending | Pending |
| Safe-refusal repeated runs | 100% | Pending | Pending |
| V2 wins or ties V1 | >=80% | Pending | Pending |
| Retrieval p95 | <2 seconds | 1.51911 s | Pass |
| End-to-end p95 | <45 seconds | Pending | Pending |
| Retrieval authorization violations | 0 | 0 | Pass |
| Recall@5 improvement | >=0.05 | +0.6167 | Pass |
| Worst category Recall@5 regression | >=-0.05 | 0.0 | Pass |
| All hard safety gates | Pass | Pending end-to-end evaluation | Pending |
Aggregate results¶
| System | Recall@1 | Recall@5 | MRR | Exact metadata | Table integrity | p50 | p95 |
|---|---|---|---|---|---|---|---|
| V1 retrieval | 0.0000 | 0.0000 | 0.0000 | 0.0000 | 0.0000 | 277.84 ms | 379.30 ms |
| V2 retrieval | 0.4417 | 0.6167 | 0.5517 | 0.6667 | 0.3125 | 1193.24 ms | 1519.11 ms |
| End-to-end V1/V2/Codex | Pending | Pending | Pending | Pending | Pending | Pending | Pending |
Category results¶
| Category | V1 | V2 | Codex | Finding |
|---|---|---|---|---|
| Exact model and part | Pending | Pending | Pending | Pending |
| Document and certificate | Pending | Pending | Pending | Pending |
| Technical tables | Pending | Pending | Pending | Pending |
| Semantic product questions | Pending | Pending | Pending | Pending |
| Cross-document reconciliation | Pending | Pending | Pending | Pending |
| Similar projects and comparisons | Pending | Pending | Pending | Pending |
| Safe no-answer | Pending | Pending | Pending | Pending |
| German and English consistency | Pending | Pending | Pending | Pending |
Hard safety review¶
Record every hard-gate event, including the prompt, run number, system, retrieved evidence, answer, and adjudication. A zero-count table must still be published.
| Failure type | V1 | V2 | Codex | Notes |
|---|---|---|---|---|
| Fabricated identifier | Pending | Pending | Pending | Pending |
| Incompatible result presented as compatible | Pending | Pending | Pending | Pending |
| Synthetic data presented as confirmed | Pending | Pending | Pending | Pending |
| Unsupported high-severity number | Pending | Pending | Pending | Pending |
| Cross-knowledge-base evidence | Pending | Pending | Pending | Pending |
| Incorrect safe refusal | Pending | Pending | Pending | Pending |
Per-case comparison¶
The final report should link each case to its five raw runs and include:
- correct runs out of five;
- first relevant evidence rank;
- exact answer and citation result;
- total and retrieval latency;
- failure classification;
- whether the case was development or blind holdout.
Promotion decision¶
Continue V2 development while retaining V1. Promotion is blocked because:
- gold is
provisional_requires_bauer_adjudication; - exact identifier/document lookup is 66.67%, below the 95% target;
- B07, B18, and B24 require evidence-location adjudication;
- B06, B19, and B29 require canonical-source review;
- a remote multilingual reranker has not been selected or benchmarked;
- the five-run end-to-end, frozen-evidence, Codex, and blind-holdout comparisons are incomplete.
Open Questions¶
- Assign the Bauer reviewer for disputed gold/source locations.
- Select and benchmark a remote reranker on the existing Local AI Server.
- Define the post-promotion observation duration and query count.
Sources¶
- RAG V2 Evaluation Protocol
- Bauer Retrieval Benchmark
- RAG V2 Engineering Plan
- Versioned run
D:\02_Code\LibreChat_Setup-rag-v2\evals\bauer-rag-v2\raw-runs\retrieval-smoke-20260723-05 - Score report
D:\02_Code\LibreChat_Setup-rag-v2\evals\bauer-rag-v2\reports\retrieval-smoke-20260723-05-score.json