Mohamed A M Elansary, PhD
Target: Senior Research Scientist, Model Evaluation — Cohere
Sourced insights (≤2)
- Claim-level RAG generation eval: Cohere treats grounded generation as too hard for binary good/bad labels — decompose into verifiable claims, then score Faithfulness (supported by retrieved docs), Correctness (vs gold), and Coverage (gold claims in the response); avoid self-judging with the generator model. Source: docs.cohere.com/page/rag-evaluation-deep-dive
- Multi-judge LLM-as-a-judge: Cohere’s retrieval-eval cookbook uses multiple Command/Aya models as independent judges with majority voting; Cohere rerank (Engine A) beat baseline Wikipedia top-n (Engine B) on 80% of queries — trustworthy, repeatable eval tooling. Source: docs.cohere.com/page/retrieval-eval-pydantic-ai
Proof — PhD UQ + production LLM shipping
- PhD Environmental Engineering, TAMUK 2022: multimodel/ensemble hydrologic uncertainty quantification and reduction on HPC (MODFLOW, VIC, PIHM, NASA LIS; USGS/NOAA/NASA) — hard-eval rigor: metric–capability alignment, failure modes, publication-grade QA.
- Production multi-tenant agentic LLM systems (GPT, Claude, Gemini): retrieval, query routing, per-tenant isolation — WhatsApp AI receptionist + voice agents in live use (Vertexium); hours reviewing complex LLM outputs for data quality.
- Strong software engineering in Python/Bash/API — prototypes that become reusable tools, not demos.
- Honest frame: eval scientist who can build — not inventing LLM-as-judge pubs or Cohere-scale eval-infra ownership.
- Brand: the PhD who ships. Prefer take-home / work-sample when interview format allows. Remote OK · EAD.