LLM evaluation · Scientific reasoning QA · Multilingual annotation

LLM evaluator for STEM reasoning, multilingual AI data, and model-output QA.

Physics-trained reviewer with experience evaluating model answers for numeric accuracy, conceptual correctness, assumption handling, and Turkish-English-Spanish language quality.

Evaluation systems for models that need more than fluent answers.

01 Gold-answer benchmark design
02 Failure-mode taxonomy
03 Human-reviewed model QA
04 Turkish / English / Spanish evaluation

Hire me for

AI model response evaluation
STEM and physics answer review
Turkish-English-Spanish AI data QA
Rubric-based annotation and labeling
Gold-answer comparison
Hallucination and reasoning-error detection
Dataset review and cleanup

Freelance services

Practical evaluation services

I support AI training and evaluation projects with careful model-output review, rubric-based labeling, STEM reasoning checks, and multilingual QA.

Model response evaluation

Reviewing AI answers for correctness, completeness, instruction following, and clarity.

STEM and physics answer review

Checking calculations, formulas, assumptions, and conceptual explanations in physics/math-style tasks.

Multilingual AI data QA

Reviewing Turkish, English, and Spanish outputs for meaning, terminology, fluency, and reasoning consistency.

Rubric-based annotation

Applying client-provided labels, grading criteria, and quality rules consistently across model outputs.

Gold-answer comparison

Comparing model responses against reference answers and identifying missing, incorrect, or unsupported claims.

Error and hallucination tagging

Marking numerical mistakes, conceptual errors, unsupported statements, prompt misreadings, and formatting failures.

Dataset cleanup and review

Checking prompt/answer pairs, improving clarity, flagging unusable data, and preparing reviewable outputs.

Work style

  • Comfortable with detailed annotation guidelines
  • Careful with edge cases and ambiguous outputs
  • Able to explain labels with short reviewer notes
  • Experienced with human review after automatic evaluation
  • Works in English, Turkish, and Spanish

Selected work

Project portfolio

A compact portfolio of published and incoming evaluation projects. The focus is evidence: what was tested, what failed, and how the results were reviewed.

Clients and work contexts

Relevant evaluation and annotation work

Client names are included where appropriate. Private platform work is described by task type rather than exposing confidential details.

Data annotation

Datamundi

Multilingual annotation and quality review involving linguistic judgment and task-specific evaluation.

AI evaluation

Mindrift

Physics and natural-science model evaluation context, including domain-specific reasoning checks.

Language services

Translation and terminology

Professional Turkish-English-Spanish language work, terminology review, and technical communication.

Private / platform work

LLM annotation clients

Human-review workflows for correctness, rubric labels, response quality, and failure-mode documentation.

Services

Evaluation work I can support

Benchmark design

Controlled prompts, perturbation variants, answer keys, and scoring rubrics.

STEM answer QA

Physics, math, and technical reasoning review for model outputs.

Failure taxonomy

Error-label systems that separate conceptual failures, numeric mistakes, and artifact issues.

Multilingual evaluation

Turkish-English-Spanish review, terminology consistency, and translation-sensitive evaluation.

Human review workflows

Annotation guidelines, review queues, disagreement handling, and final cleaned outputs.

Dataset packaging

CSV, JSONL, Excel, Markdown, GitHub README, and report-ready deliverables.

Background

Physics-trained evaluator focused on quantum reasoning, STEM QA, and multilingual model evaluation.

My academic background is in physics, with a current focus on quantum-computation reasoning, scientific problem evaluation, and the failure modes of language models in technical domains. I work on evaluation projects where the important question is not only whether a model gives the right final answer, but whether it preserves the correct assumptions, cost model, notation, and reasoning structure.

My current portfolio connects physics reasoning with LLM evaluation: partial quantum search, query-complexity accounting, hallucination stress tests, Turkish scientific reasoning, and trilingual benchmark design. The goal is to turn model failures into inspectable evidence: gold answers, error labels, human-reviewed notes, and reusable datasets.

Core area
LLM evaluation for quantum computation, physics, and STEM reasoning
Languages
Turkish, English, Spanish
Location
📍 Barcelona, Spain

Contact

Need a benchmark, dataset, or evaluation workflow?

Send the domain, model type, language, and what you need to measure.