Skip to content
AI & Automation

AI Data Annotation: How to Prepare High-Quality Datasets

How to run data annotation for AI: label taxonomies, annotation guidelines, workflows and tooling, model-assisted labelling, quality control, inter-annotator agreement, expert review, dataset documentation and versioning.

Quick answer

High-quality annotation starts with a taxonomy that matches the decisions your AI must make and guidelines with definitions, examples and edge-case rules. Pilot with a small batch, measure agreement between annotators, refine guidelines, then scale with model-assisted pre-labelling, review and adjudication by experts. Use gold-standard items and audits to monitor quality, document the dataset and version it alongside the models and prompts it supports.

Where This Fits

Image-specific labelling is covered in AI image recognition, dataset checks in data quality for AI, generated alternatives in synthetic data generation and how labelled sets drive testing in LLM evaluation pipeline.

What Gets Annotated

DataAnnotation typesTypical use
TextCategories, entities, sentiment, relevance, spansClassification, extraction, search evaluation
ImagesClasses, bounding boxes, segmentation masksRecognition, inspection
AudioTranscripts, speakers, eventsSpeech systems, call analytics
DocumentsField values, layout regions, tablesExtraction evaluation and training
Model outputsRatings, preferences, error categoriesLLM evaluation, fine-tuning

Designing the Taxonomy and Guidelines

Most annotation quality problems start with the taxonomy. Classes should map to decisions the system must make, not to every distinction someone can imagine. Define each class in plain language, give positive and negative examples, explain how to handle overlaps and include 'other' and 'unclear' options so annotators are not forced into wrong labels.

Guidelines evolve. Version them, record why each change was made and re-label affected items when definitions change materially. Annotators' questions are the best source of improvements.

The Annotation Workflow

Review and adjudication turn individual judgements into a dataset you can trust.

Need labelled data for an AI project?

ZSpace Labs designs annotation programmes, guidelines and quality checks for AI teams. See AI development services.

Start a Project

Measuring Quality

Use several controls together. Agreement: have multiple annotators label an overlapping subset and compute agreement; investigate low-agreement classes. Gold items: mix in items with known correct labels to measure each annotator's accuracy. Review: a second person checks a sample of every annotator's work. Adjudication: an expert resolves disagreements and records the decision as guidance.

Low agreement is information, not just failure. It may mean guidelines are unclear, classes overlap or the task is genuinely ambiguous, in which case the AI system should probably express uncertainty too.

Model-Assisted Labelling

Pre-labelling with an existing model, or with an LLM for text, can multiply annotator throughput: people confirm or correct rather than starting from scratch. The risk is anchoring, where annotators accept plausible but wrong suggestions. Audit pre-labelled items, track how often annotators change suggestions, and periodically label a sample without suggestions to compare. Active learning, where the model selects uncertain items for labelling, focuses effort where it adds most.

Tooling and Data Security

Tools such as Label Studio and CVAT support many data types, workflows and exports; commercial platforms add workforce management and analytics. Evaluate data security carefully, especially when using external annotators or vendors: access controls, where data is stored, whether annotators can download data and how personal information is handled. Redact or pseudonymize where the task allows.

Documentation and Versioning

Document each dataset with its purpose, sources, collection period, taxonomy and guideline versions, annotator profile, agreement scores, known gaps and splits. The datasheets for datasets proposal is a practical template. Version datasets so each model or prompt evaluation can be tied to the exact data used; see AI data lineage.

Advantages and Limitations

Well-run annotation produces datasets that make models and evaluations trustworthy. It is labour-intensive, specialised labels are expensive and quality drifts without ongoing controls. Spend annotation effort where errors are costly and where models struggle, rather than labelling everything uniformly.

How to Run an Annotation Project Step by Step

  • 1. Define decisions and taxonomy with domain owners
  • 2. Write guidelines with examples and edge-case rules
  • 3. Pilot 100 to 200 items with several annotators and measure agreement
  • 4. Refine guidelines and repeat until agreement is acceptable
  • 5. Scale with pre-labelling, review and gold items
  • 6. Adjudicate disagreements and feed decisions into guidelines
  • 7. Document and version every release of the dataset

Annotating LLM Outputs and Preferences

Generative AI introduces a new annotation task: judging model outputs. Reviewers rate answers against rubrics, compare two outputs side by side, categorize errors (factual, incomplete, unsafe, off-topic) or write corrected versions. These labels calibrate automated judges, build evaluation sets and, where used, provide preference data for fine-tuning. Rubrics need the same care as classification taxonomies: clear criteria, examples of each score and agreement checks. Blind reviewers to which model or version produced each output. See AI model evaluation.

Working With Annotation Vendors

External annotation providers add capacity but need careful management. Share guidelines and gold items, run a paid pilot and measure agreement before scaling, require data security terms covering storage, access, retention and subcontracting, and keep an internal expert reviewing samples continuously. Agree how disagreements and guideline questions are escalated. For sensitive data, prefer redacted or synthetic inputs where the task allows, or keep annotation in-house.

Example Guideline Entry

Guidelines work best as a set of short entries, one per class or decision, each with a definition, examples and the rules annotators actually need for borderline cases.

Example: annotation guideline entry (illustrative)
class: billing_dispute
definition: Customer disagrees with a charge already made.
include:
  - "I was charged twice for March"
  - "This invoice includes a seat we cancelled"
exclude:
  - Questions about how billing works (-> billing_question)
  - Requests to change plan (-> plan_change)
edge_cases:
  - Dispute + cancellation request: label billing_dispute (primary issue rule)
  - Unclear if charged yet: label billing_question, flag 'unclear'
version: 4 (2026-09-30) - added primary issue rule after low agreement

Worked Example

An illustrative scenario, not a client case: a support team labels tickets into 25 categories for a routing model, but agreement between annotators is low. Analysis shows three pairs of overlapping categories and no rule for multi-issue tickets. The team merges overlaps, adds a primary-issue rule and examples, and agreement on the pilot set improves enough to scale labelling with model pre-labels and weekly audits.

Common Mistakes

  • Taxonomies with overlapping or vague classes
  • Scaling before agreement is measured
  • Accepting pre-labels without audits
  • No version control for guidelines or datasets
  • Using non-experts for specialised judgements without expert review

Want a quality review of your training or evaluation data?

Talk to ZSpace Labs about annotation and dataset quality for your AI systems.

Start a Project

Conclusion

Labels are only as good as the taxonomy, guidelines and quality controls behind them. Pilot, measure agreement, refine, then scale with review and documentation.

FAQ

Common questions

Adding labels or structured information to raw data, such as categories for text, bounding boxes on images, transcripts for audio or quality ratings for model outputs, so the data can be used to train or evaluate AI systems.

Related services
Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.