Skip to content

V1, V2, and Codex Evaluation Results

Status: Development retrieval smoke complete; formal comparison not started V1 baseline: Available in the existing Bauer Retrieval Benchmark V2 result: Retrieval improves, but exact lookup misses the provisional promotion gate Promotion decision: Continue private V2 evaluation; retain normal Bauer on V1

These are measured development retrieval results produced under the RAG V2 Evaluation Protocol. They validate the deployed private path but do not substitute for the required verified-gold, five-run, end-to-end, frozen-evidence, and blind-holdout comparison.

Run identity

Item Value
Evaluation date 2026-07-23
Run retrieval-smoke-20260723-05; 30 development cases, one V1/V2 repetition
Corpus manifest checksum 2934bb3bf75514a4ae4d340cb68c023831ca20a5e89e194fd20774b1e561b0d8
V1 source baseline 2815a8f2ba1b7db04d40a80934544f092e54518b
V2 evaluated source bf3c3cb0a96f054ad064824623baf25b4785005b
RAG / LibreChat deployments 4c779152-def4-4c9a-a245-3e205a603547 / 1142dfd1-fc50-477a-8cc4-7f226cc212c2
V1 / V2 Agent IDs agent_Z8A2LtQWLeP4KuUDbvqZL / agent_pmPMcA25UXS7vznUaz-DU
V2 index bauer-rag-v2-2026-07; run 7e23f6e1-d876-4da0-a677-9a3cdb37db36
Embedding local/embed-engineering-1024
Reranker deterministic-multilingual-fallback-v1; remote not configured
Answer model local/qwen-coder; not exercised by this retrieval-only smoke
Evaluation artifact commit d25468e
Raw-run / score SHA-256 fdaf707c0ddf5000712ef176b583b07ae2030fc110e4a0cbe2579177ac534d4a / b929e17c24a7955cd8ec0457898d22a6121e5cf6420aa877df24281e62330a5b

Promotion gates

Gate Required Measured Result
Exact identifier/document lookup >=95% 66.67% Fail against provisional locations
Exact-answer accuracy >=90% Pending Pending
Citation correctness >=95% Pending Pending
Safe-refusal repeated runs 100% Pending Pending
V2 wins or ties V1 >=80% Pending Pending
Retrieval p95 <2 seconds 1.51911 s Pass
End-to-end p95 <45 seconds Pending Pending
Retrieval authorization violations 0 0 Pass
Recall@5 improvement >=0.05 +0.6167 Pass
Worst category Recall@5 regression >=-0.05 0.0 Pass
All hard safety gates Pass Pending end-to-end evaluation Pending

Aggregate results

System Recall@1 Recall@5 MRR Exact metadata Table integrity p50 p95
V1 retrieval 0.0000 0.0000 0.0000 0.0000 0.0000 277.84 ms 379.30 ms
V2 retrieval 0.4417 0.6167 0.5517 0.6667 0.3125 1193.24 ms 1519.11 ms
End-to-end V1/V2/Codex Pending Pending Pending Pending Pending Pending Pending

Category results

Category V1 V2 Codex Finding
Exact model and part Pending Pending Pending Pending
Document and certificate Pending Pending Pending Pending
Technical tables Pending Pending Pending Pending
Semantic product questions Pending Pending Pending Pending
Cross-document reconciliation Pending Pending Pending Pending
Similar projects and comparisons Pending Pending Pending Pending
Safe no-answer Pending Pending Pending Pending
German and English consistency Pending Pending Pending Pending

Hard safety review

Record every hard-gate event, including the prompt, run number, system, retrieved evidence, answer, and adjudication. A zero-count table must still be published.

Failure type V1 V2 Codex Notes
Fabricated identifier Pending Pending Pending Pending
Incompatible result presented as compatible Pending Pending Pending Pending
Synthetic data presented as confirmed Pending Pending Pending Pending
Unsupported high-severity number Pending Pending Pending Pending
Cross-knowledge-base evidence Pending Pending Pending Pending
Incorrect safe refusal Pending Pending Pending Pending

Per-case comparison

The final report should link each case to its five raw runs and include:

  • correct runs out of five;
  • first relevant evidence rank;
  • exact answer and citation result;
  • total and retrieval latency;
  • failure classification;
  • whether the case was development or blind holdout.

Promotion decision

Continue V2 development while retaining V1. Promotion is blocked because:

  • gold is provisional_requires_bauer_adjudication;
  • exact identifier/document lookup is 66.67%, below the 95% target;
  • B07, B18, and B24 require evidence-location adjudication;
  • B06, B19, and B29 require canonical-source review;
  • a remote multilingual reranker has not been selected or benchmarked;
  • the five-run end-to-end, frozen-evidence, Codex, and blind-holdout comparisons are incomplete.

Open Questions

  • Assign the Bauer reviewer for disputed gold/source locations.
  • Select and benchmark a remote reranker on the existing Local AI Server.
  • Define the post-promotion observation duration and query count.

Sources