Model response evaluation
Reviewing AI answers for correctness, completeness, instruction following, and clarity.
LLM evaluation · Scientific reasoning QA · Multilingual annotation
Physics-trained reviewer with experience evaluating model answers for numeric accuracy, conceptual correctness, assumption handling, and Turkish-English-Spanish language quality.
Evaluation systems for models that need more than fluent answers.
Freelance services
I support AI training and evaluation projects with careful model-output review, rubric-based labeling, STEM reasoning checks, and multilingual QA.
Reviewing AI answers for correctness, completeness, instruction following, and clarity.
Checking calculations, formulas, assumptions, and conceptual explanations in physics/math-style tasks.
Reviewing Turkish, English, and Spanish outputs for meaning, terminology, fluency, and reasoning consistency.
Applying client-provided labels, grading criteria, and quality rules consistently across model outputs.
Comparing model responses against reference answers and identifying missing, incorrect, or unsupported claims.
Marking numerical mistakes, conceptual errors, unsupported statements, prompt misreadings, and formatting failures.
Checking prompt/answer pairs, improving clarity, flagging unusable data, and preparing reviewable outputs.
Selected work
A compact portfolio of published and incoming evaluation projects. The focus is evidence: what was tested, what failed, and how the results were reviewed.
In this project, I built and reviewed a two-stage benchmark for quantum-search reasoning. The goal was to check whether model answers preserved the correct math, assumptions, and reasoning structure under initial prompts and targeted follow-up questions.
The same evaluation workflow can be adapted to STEM answer review, multilingual AI data QA, rubric-based annotation, and model-output quality checks.
A local-first evaluation pipeline for testing whether language models can reason through Partial Quantum Search problems beyond plausible final answers.
Stage 1 evaluates first-answer correctness. Stage 2 tests whether the same reasoning survives targeted follow-up pressure.
Of 112 successful outputs, 58 failed review (52%) and 19 of those were critical. Figures match the published Stage 1 report (original 120-run set).
By branch — Robustness 153 (114 P / 20 partial / 19 F) · Conceptual correction 24 (9 P / 15 F, 37.5% pass) · Numeric audit 12 (4 P / 2 partial / 6 F).
Compared deterministic scientific grading against BERTScore, ROUGE-L, chrF, and ROSCOE across 252 structured LLM outputs. The dashboard shows that text/reasoning metrics are useful diagnostics but weak correctness proxies, and that large numeric failures require deterministic gates.
Designed Partial Quantum Search scenarios to test scientific reasoning, assumption tracking, numerical consistency, and cost-model interpretation.
Collected model answers to controlled PQS prompts and checked whether models produced correct reasoning, not only plausible final numbers.
Used deterministic extraction, numeric checks, concept checklists, and human review to classify answers into pass, numeric error, critical error, artifact, or unusable cases.
Generated follow-up prompts from Stage 1 outcomes: robustness checks for passes, numeric audits for arithmetic errors, and conceptual correction prompts for critical failures.
Exported reviewed evidence, aggregate results, and report-ready artifacts for a transparent portfolio case study.
Clients and work contexts
Client names are included where appropriate. Private platform work is described by task type rather than exposing confidential details.
Multilingual annotation and quality review involving linguistic judgment and task-specific evaluation.
Physics and natural-science model evaluation context, including domain-specific reasoning checks.
Professional Turkish-English-Spanish language work, terminology review, and technical communication.
Human-review workflows for correctness, rubric labels, response quality, and failure-mode documentation.
Services
Controlled prompts, perturbation variants, answer keys, and scoring rubrics.
Physics, math, and technical reasoning review for model outputs.
Error-label systems that separate conceptual failures, numeric mistakes, and artifact issues.
Turkish-English-Spanish review, terminology consistency, and translation-sensitive evaluation.
Annotation guidelines, review queues, disagreement handling, and final cleaned outputs.
CSV, JSONL, Excel, Markdown, GitHub README, and report-ready deliverables.
Background
My academic background is in physics, with a current focus on quantum-computation reasoning, scientific problem evaluation, and the failure modes of language models in technical domains. I work on evaluation projects where the important question is not only whether a model gives the right final answer, but whether it preserves the correct assumptions, cost model, notation, and reasoning structure.
My current portfolio connects physics reasoning with LLM evaluation: partial quantum search, query-complexity accounting, hallucination stress tests, Turkish scientific reasoning, and trilingual benchmark design. The goal is to turn model failures into inspectable evidence: gold answers, error labels, human-reviewed notes, and reusable datasets.
Contact
Send the domain, model type, language, and what you need to measure.