AI work directory Red Pen Jobs

How to get started with AI training work

AI training work is contract work where you help improve artificial intelligence models. You might write prompts, evaluate the quality of AI-generated answers, label images, annotate data, or verify that a model’s output is factually correct. The work is remote, it is usually paid hourly or per task, and it is available to professionals in dozens of fields.

What the work looks like

The common thread is human judgement. A model generates an output — a paragraph of text, a code snippet, a translation, a summary — and a human decides whether it is good, bad or somewhere in between.

Tasks vary by role:

  • Prompt writing: you write questions or instructions that test what a model can do
  • Response evaluation: you read two or more AI-generated answers and rank them by quality, accuracy or helpfulness
  • Data labelling: you tag images, text or video with categories the model needs to learn
  • Domain review: you apply professional expertise (law, medicine, engineering, finance) to verify that a model’s output is correct in your field
  • Red-teaming: you try to make a model produce incorrect, harmful or policy-violating responses so its weaknesses can be fixed

Most tasks take between two and thirty minutes. You work when you choose, from wherever you are, on a laptop with a browser.

Where to find the work

This site tracks roles from four platforms:

  • micro1 — the largest by openings, covering everything from generalist annotation to specialist expert review
  • Mercor — skews toward professional and technical roles: software engineers, data scientists, subject-matter experts
  • Terac — often posts short-form evaluation tasks and interviews, many paid per session
  • AfterQuery — smaller catalogue, narrower skill focus

You can browse all open roles or filter by your profession. Each listing shows the pay the platform states, the number of openings, and whether it is restricted to certain countries.

What the screening process is like

Every platform screens applicants. What that looks like varies.

On micro1, you complete a skills assessment relevant to the role — a coding test for engineering roles, a writing sample for content roles, a domain quiz for specialist positions. The process can take thirty minutes to a few hours depending on the role.

On Mercor, screening often includes a technical assessment and may involve a short interview or work sample. Mercor tends to be more selective for its higher-paying roles.

On Terac, some roles begin with the evaluation itself: the screening is the task, and you are paid for completing it.

There is no single “AI training certification” that gets you in. Each platform and each role runs its own screening, and what matters is whether you can do the specific work the role asks for.

What you need

For most roles:

  • A laptop or desktop with a reliable internet connection
  • A browser (most platforms run entirely in the browser)
  • Fluency in the language the role specifies (English for most, but many roles need other languages)
  • The domain knowledge the role asks for — if a role says “physics PhD required”, it means it

For generalist roles (data labelling, basic evaluation), the bar is lower: clear thinking, attention to detail, and following annotation guidelines consistently.

For specialist roles (legal review, medical evaluation, advanced coding), you need the credentials and experience the posting describes. These roles pay more because fewer people qualify.

What to expect on your first day

You will be given a set of guidelines — often a detailed document explaining exactly how to evaluate, label or annotate. Read it carefully. The guidelines are the job. Consistency with the guidelines matters more than your personal opinion about what a “good” answer looks like.

Most platforms start you with a small batch of tasks to check your work quality. If your accuracy is high, you get access to more work. If it is low, you get feedback or are moved to a different task type.

The work is flexible: you log in when you want, pick up available tasks, and stop when you are done. But “flexible” does not mean “guaranteed”. Task availability fluctuates, and a role listing 10,000 openings does not mean 10,000 hours of work are waiting for you personally.

Next steps

  1. Browse open roles to see what matches your skills
  2. Read about what AI training work pays to see the rate range
  3. Check the platform pages for details on each platform’s screening and terms