Amazon + Outlier · AI evaluation
Applying editorial judgmentto AI evaluation.
How I use structured criteria, comparative judgment, written rationales, and targeted rewrites to identify why generated responses succeed or fail.
The work
Editorial judgment made explicit
A strong evaluation does more than label one response “better.” It identifies the requirements in the prompt, tests each response against consistent criteria, explains the tradeoffs, and shows how the weaker output could be improved.
At Amazon AGI Data Services, I evaluate generative, conversational, and agentic outputs against structured standards; document decisions and recurring quality patterns; and contribute to evaluation frameworks, prompt guidance, and process documentation for human-in-the-loop work.
As an Outlier contributor, I evaluate and rank generated responses for instruction following, factual accuracy, relevance, reasoning, voice and tone, and safety. I write detailed rationales, rewrite weak outputs, and review voice-assistant responses and transcripts.
The criteria
Define quality for the task
The criteria change by task. The discipline is to apply the selected criteria consistently and explain the decision precisely enough that another reviewer can follow it.
Reconstructed example
Make the decision traceable
This example is fictional. It demonstrates the method without reproducing client prompts, outputs, interfaces, or data.
Prompt
Write a response to a customer whose online payment was declined. Use 55 words or fewer. Acknowledge the issue, do not speculate about the cause, and give exactly one next step: contact the card issuer. Do not promise resolution.
Candidate A
Your bank probably declined the charge because it looked suspicious or your balance was too low. Check your balance, try the payment again, or use another card. If that still does not work, call your bank and they should be able to fix it.
Speculates · Adds steps · Promises an outcome
Candidate B
I’m sorry your payment didn’t go through. Please contact your card issuer for more information about the declined charge.
Acknowledges · Complies · Stays within scope
Ranking and rationale
Candidate B is stronger. It acknowledges the issue, stays within the requested length, does not guess why the payment was declined, gives exactly one next step, and avoids promising resolution. Candidate A violates several explicit constraints.
Targeted rewrite
I’m sorry your payment didn’t go through. Please contact your card issuer to ask why the charge was declined and what you need to do next.
The rewrite preserves Candidate B’s compliance while making the purpose of the contact more useful.
Voice evaluation
Separate content from delivery
Voice work adds another layer: the text may be acceptable while the spoken delivery fails the intended persona or interaction. My review includes creating or correcting transcripts and evaluating whether tone, accent, pacing, delivery, and overall conversational quality match the system prompt and rubric.
The same discipline applies: identify the specific failure and support the rating with observable evidence.
What it demonstrates
Quality control under ambiguity
- Consistent use of task-specific rubrics
- Precise explanations for comparative rankings
- Targeted rewriting rather than generic criticism
- Recognition of recurring quality patterns
- Quality control for human-in-the-loop AI systems
These are demonstrated capabilities—not a claim that my individual evaluation caused a measurable improvement in a commercial model.
The most useful feedback names the smallest meaningful failure.