How Training Data
Is Collected, and
Who Collects It
Who taught your phone to understand your accent, read your handwriting, and recognize your street? Thousands of people who contributed theirs, and they were paid for it. That work has a name.
Speech, photos, handwriting and video · Defined projects · Paid per approved unit
Experts accepted daily
Expert domains
Countries
Data collection is the gathering of real-world examples that technology learns from: speech, photos, video, handwriting and text. In Human Intelligence work it means paid, structured projects where people contribute those examples deliberately, following written guidelines they agree to before starting.
Every contribution is reviewed, and approved contributions become the dataset behind a feature people use. It happens in more than 300 languages, on ordinary phones, from home.
If you speak a language and carry a phone, you are already qualified to do it.
Why technology cannot invent its own examples
Training data does not exist until a person contributes it.
Technology can generate endless artificial speech, text and images. What it cannot generate is the one thing it needs most: something it has never encountered.
That is the part people find surprising. Technology is not limited by computing power. It is limited by examples: real accents in real kitchens, a street sign in a script no map has read, handwriting that slants the way yours does, the same word said differently on either side of a border, a voice a caption tool has never had to follow. None of it can be simulated into existence, because the whole point is that it came from the world.
Three gaps technology cannot close on its own
Accents it has never heard
Speech technology fails hardest for the voices it saw least. The fix is not a better algorithm. The fix is recordings from the people it keeps misunderstanding.
Places nobody photographed
Street signs, storefronts, handwriting, home interiors. A system knows only the neighborhoods its dataset came from, and everywhere else stays a blind spot until someone who lives there fills it.
The way people live at home
Studio recordings and staged photographs are clean, scripted and close to useless. Real speech overlaps and trails off. Real rooms are cluttered and badly lit. Real handwriting is rushed. Only real people produce any of it.
Why this work matters more than it sounds
Most of the world’s languages have never had enough recorded speech or written text for technology to learn from properly. The same is true of most handwriting, most street signs, and most of the ways people speak when they are at home. Everyone inside that gap is asked to use a phone, a bank, a health service or a classroom in a language the system handles badly, or in someone else’s language entirely.
That gap does not close by itself, and no amount of computing power closes it either. It closes one contribution at a time, when a speaker of a language decides to put it on the record.
This is the quiet part of how technology gets built. Not the clever part. The evidence underneath it. A screen reader that handles an accent it was never given, a camera that reads Devanagari handwriting, a car that hears a command over a noisy engine: each of those started as somebody sitting in their own home with their own phone, following a brief, contributing something only they had.
Technology inherits whoever showed up. That is the whole argument for showing up.
How the work happens
Human Intelligence work in data collection follows the same six steps across most projects. Only the first one feels unfamiliar. After that it settles into a rhythm people tend to describe as unfussy: read the brief, do the work, see what passed.
Project qualification
You confirm your language, your country and the device already in your pocket, then contribute a short sample. It runs both ways. The project checks that your contribution fits, and you find out what it will ask of you before you commit to anything. Gates differ: some projects need a specific dialect, some need a specific phone camera.
Guidelines
Every project arrives with written guidelines covering the environment, the framing, the length, the format, and what a correct contribution looks like. They are more specific than people expect, and that specificity is a courtesy. Nobody is left guessing what good means. Reading them is part of the paid time.
Collection
This is the part that is yours. You record, photograph or write to the brief, on your own device, in your own home, on your own schedule. A session might be reading prompts aloud at the kitchen table, photographing everyday objects on a walk you were taking anyway, or recording a conversation with people you know. It rarely looks like work from the outside. It is.
Quality review
Every unit is checked against the guidelines. Where something misses, you are told why in specific terms, and most projects allow a re-take. That loop is the difference between guessing and getting steadily better at it.
Approval
Approved units are what count. On many projects reviewers also compare contributions across everyone taking part, because a dataset that leans too heavily on one voice or one street teaches the wrong lesson. Your contribution is measured on its own terms and on how it balances the whole.
Payment
You are paid per approved unit, twice monthly, by PayPal or Payoneer. The rate is shown before you agree to anything.
Six steps, and only the third one needs you specifically. That is the whole design.
Where this shows up in everyday technology
Data collection is not one thing. The examples technology needs change by field, and you have met most of these already without being told what was underneath them.
Voice recognition
This is where the question how does speech recognition work leads: back to the recordings it learned from.
- Wake words and voice commands
- Dictation, captions and transcription
- Real conversation, in every accent that uses it
When a device finally understands somebody it used to talk over, that is the technology having finally heard someone like them.
Cameras and image editing
Scene detection, background removal, document scanning and image search all work from photographs of ordinary life.
- A receipt on a table
- A sign at the end of a street
- A shelf in a corner shop
Nothing staged, nothing lit. The value is in how ordinary it is.
Accessibility
Technology that reads a screen aloud, captions a room in real time, or lets somebody operate a device without touching it.
- Live captions and transcription
- Voice control and screen reading
- Speech that is atypical or impaired
The people these features exist for are the people whose data is scarcest. That is the whole gap, in one line.
Handwriting and documents
Systems that read handwriting work from people writing exactly as they normally write.
- Handwriting in any script
- Forms, notes and filled-in documents
- Personal shorthand and slant
Handwriting sits close to a fingerprint, which is why it takes thousands of samples before yours can be read.
Maps, cars and smart homes
Street-level detail and in-car voice control come from real streets and real rooms, noise included.
- Street signs and shopfronts
- Commands over a running engine
- Rooms with the television on
A recording made in a silent room teaches a device almost nothing about the room it is going into.
Translation and language tools
Parallel recordings and text keep smaller languages visible to the technology everyone else already takes for granted.
- More than 300 languages
- Matched recordings and text
- Dialects that data has skipped
A language with no data does not get a worse version of these tools. It gets none.
Different fields, one pattern. The technology stops improving exactly where the examples run out.
What we mean by an expert
Everyone is an expert.
Not a credential. Not a title. An expert is someone who knows something the technology does not, and could only have learned it by living it: a language, a dialect, a trade, a place, a way of holding a pen.
On a data collection project that expertise is your accent, your street, your handwriting, your hand. You already have it. Nobody taught it to you for this.
Who takes part
Not engineers. The projects that need the widest range of contributions need the widest range of people.
OneForma accepts more than 830 experts a day, across more than 100 countries and more than 300 languages. Data collection is where most of them started.
What it is not
Being straight about this saves everyone time.
It is not taken from you
Nothing here is gathered in the background. Every project states what is being collected and how it will be used before you agree, and you contribute deliberately, item by item.
These are projects, not a position
You take the ones you want, and there is nothing owed between them.
It is not a course
Nobody is teaching you to speak your language or photograph your street. What you already have is the whole contribution.
It is not always exciting
Good contributions follow the brief exactly, and the brief can be picky. The people who enjoy this tend to like precise, finishable work.
Questions people ask.
What is data collection?
Data collection is the gathering of real-world examples that technology learns from: speech, photos, video, handwriting and text. In Human Intelligence work it means paid, structured projects where people contribute those examples deliberately, following written guidelines, under terms they agree to first.
How is training data collected?
Through defined projects. Contributors qualify for a project, read its guidelines, then record, photograph or write to a brief on their own devices. Every contribution is reviewed against the guidelines, and approved units become the dataset behind the finished feature.
What is training data?
Training data is the set of real examples a system learns patterns from. For voice recognition it is recorded voices; for photo and editing tools it is photographs; for handwriting recognition it is handwritten samples. The range of the data sets the limits of what the finished technology can handle.
Why can’t training data be generated automatically?
Because a system can only recombine what it has already seen. It cannot produce an accent, a script or a street it has no examples of. Anything genuinely new to the system has to come from the world, which means it has to come from a person.
Do I need special equipment?
Almost never. Most projects are built for the phone you already own, because that is the device the technology has to work on. Where a project needs specific hardware, it says so before you qualify.
Do I need experience with technology?
No. You need your language, your voice or your camera, and the patience to follow guidelines exactly. The technical side is handled by the project.
What happens to what I contribute?
Each project states its use before you join: what is collected, what it builds, and the terms that cover it. Your contribution becomes part of the dataset for that project. Nothing is collected outside what the project describes.
What does OneForma mean by an expert?
Everyone is an expert. Not a credential and not a title: an expert is someone who knows something the technology does not, and could only have learned it by living it. A language, a dialect, a trade, a place, a way of holding a pen. On a data collection project that expertise is your accent, your street or your handwriting.
What does it pay?
Rates are set per project and shown before you agree to anything. You are paid per approved unit, twice monthly, by PayPal or Payoneer.
The technology being built now knows only what people chose to contribute. Your accent, your street, your handwriting: it has probably met none of them.