Amazon + Outlier · AI evaluation

Applying editorial judgmentto AI evaluation.

How I use structured criteria, comparative judgment, written rationales, and targeted rewrites to identify why generated responses succeed or fail.

OrganizationsAmazon AGI Data Services + Outlier
RolesAI Content Specialist; AI Model Trainer
Timeframe2024–present
ScopeEvaluation, ranking, rationales, rewrites
Additional workPattern analysis, prompt guidance, voice review
EvidenceReconstructed public example

The work

Editorial judgment made explicit

A strong evaluation does more than label one response “better.” It identifies the requirements in the prompt, tests each response against consistent criteria, explains the tradeoffs, and shows how the weaker output could be improved.

At Amazon AGI Data Services, I evaluate generative, conversational, and agentic outputs against structured standards; document decisions and recurring quality patterns; and contribute to evaluation frameworks, prompt guidance, and process documentation for human-in-the-loop work.

As an Outlier contributor, I evaluate and rank generated responses for instruction following, factual accuracy, relevance, reasoning, voice and tone, and safety. I write detailed rationales, rewrite weak outputs, and review voice-assistant responses and transcripts.

The criteria

Define quality for the task

Instruction followingDid the response satisfy every explicit constraint?
AccuracyAre its claims supportable and appropriately qualified?
RelevanceDoes it address the actual need without wandering?
ReasoningDoes the conclusion follow from the information available?
ClarityIs the response organized, readable, and specific?
Voice + toneDoes the language fit the audience and situation?
Responsible behaviorDoes it avoid unsafe guidance or unsupported certainty?

The criteria change by task. The discipline is to apply the selected criteria consistently and explain the decision precisely enough that another reviewer can follow it.

Reconstructed example

Make the decision traceable

This example is fictional. It demonstrates the method without reproducing client prompts, outputs, interfaces, or data.

Prompt

Write a response to a customer whose online payment was declined. Use 55 words or fewer. Acknowledge the issue, do not speculate about the cause, and give exactly one next step: contact the card issuer. Do not promise resolution.

Candidate A

Your bank probably declined the charge because it looked suspicious or your balance was too low. Check your balance, try the payment again, or use another card. If that still does not work, call your bank and they should be able to fix it.

Speculates · Adds steps · Promises an outcome

Candidate B

I’m sorry your payment didn’t go through. Please contact your card issuer for more information about the declined charge.

Acknowledges · Complies · Stays within scope

Ranking and rationale

Candidate B is stronger. It acknowledges the issue, stays within the requested length, does not guess why the payment was declined, gives exactly one next step, and avoids promising resolution. Candidate A violates several explicit constraints.

Targeted rewrite

I’m sorry your payment didn’t go through. Please contact your card issuer to ask why the charge was declined and what you need to do next.

The rewrite preserves Candidate B’s compliance while making the purpose of the contact more useful.

Voice evaluation

Separate content from delivery

Voice work adds another layer: the text may be acceptable while the spoken delivery fails the intended persona or interaction. My review includes creating or correcting transcripts and evaluating whether tone, accent, pacing, delivery, and overall conversational quality match the system prompt and rubric.

The same discipline applies: identify the specific failure and support the rating with observable evidence.

What it demonstrates

Quality control under ambiguity

  • Consistent use of task-specific rubrics
  • Precise explanations for comparative rankings
  • Targeted rewriting rather than generic criticism
  • Recognition of recurring quality patterns
  • Quality control for human-in-the-loop AI systems

These are demonstrated capabilities—not a claim that my individual evaluation caused a measurable improvement in a commercial model.

The most useful feedback names the smallest meaningful failure.