Intraoral dataset

18,000 intraoral photographs, annotated at pixel level

Expert segmentation for dental AI: six clinical classes, six standardised views per patient, drawn from a clinical pipeline in production.

  • Tooth
  • Gum
  • Tartar
  • Cavity
  • Gingivitis
  • Restoration
18,000+

intraoral photographs

~3,000

unique patients

100,000+

polygon instances

480 images across 80 patients, published under approved access — masks, COCO export and data card included.

Overview

What the dataset is

A curated collection of intraoral photographs carrying pixel-level polygon annotations, produced by Dr Romain Lateur, dental surgeon, and drawn from a clinical pipeline running in dental practices today.

Built for semantic segmentation: a model learns to identify and localise oral structures and pathologies directly on standard dental photographs — not radiographs.

Example opposite: P2310, upper arch, tartar.

Annotated intraoral photograph — upper arch, tartar
Annotated classes

Six clinical classes, anatomy and pathology

Colours match the exported mask channels and the on-image legend exactly.

Tooth

tooth

Visible enamel surface — the most frequent class by pixel area.

Gum

gum

Soft gingival tissue around the teeth. Colour varies across individuals.

Tartar

tartar

Hardened plaque deposits — small, irregular, periodontally relevant.

Cavity

cavity

Decayed tooth regions. Darker, structurally irregular, often subtle.

Gingivitis

gingivitis

Gum inflammation: redness and swelling, recognised by texture and colour.

Restoration

restoration

Crowns and restorative material — distinct colour and reflectivity.

The capture protocol

Six standardised views per patient

One complete patient (P2389), the six-view series in capture order. Every patient is photographed through the same six positions, in the same order — so view, arch and sextant are always known, never inferred.

  1. 1upper_center view of patient P2389, annotated

    upper_center

  2. 2upper_right view of patient P2389, annotated

    upper_right

  3. 3upper_left view of patient P2389, annotated

    upper_left

  4. 4lower_center view of patient P2389, annotated

    lower_center

  5. 5lower_right view of patient P2389, annotated

    lower_right

  6. 6lower_left view of patient P2389, annotated

    lower_left

Annotation

Multi-class semantic segmentation

Closed polygons, not bounding boxes — the true boundary of every lesion is traced.

annotation type
Closed polygons (2D coordinate sequences)
storage
One JSON file per image, plus a standard COCO export
masks
Multi-channel binary, shape (6, H, W) — one channel per class
overlap
Separate channel per class: overlapping findings are preserved (tartar on tooth)
localisation
Every finding is tied to the tooth, the gum, or the tooth-gum junction
P2364, upper arch — all six classes visible in a single frame
P2364 · upper_center · all six classes visible in a single frame
Annotated examples

Photograph, polygons and legend — delivered together

Every image in the dataset ships with its polygons and its legend, in the same export.

P2364 · lower_center — tartar & gingivitis
P2364 · lower_centertartar & gingivitis
P2356 · lower_center — gingivitis, clinical setting
P2356 · lower_centergingivitis, clinical setting
P804 · upper_center — heavy tartar
P804 · upper_centerheavy tartar
P6 · upper_center — cavity
P6 · upper_centercavity
Scale & class balance

What 18,000 images really contain

Two anatomical classes in almost every frame, four pathologies at their real-world prevalence. Nothing is rebalanced or resampled: this is the imbalance a model meets in practice.

ClassInstancesImagesPatientsPer positive image
Toothanatomical base30.3 %~100 %100 %2.1
Gumanatomical base41.3 %~98 %100 %2.9
Tartarpathology19.6 %~33 %~54 %4.0
Cavitypathology6.4 %~25 %~49 %1.8
Gingivitispathology0.9 %~5 %~14 %1.4
Restorationpathology1.4 %~6 %~12 %1.5

100,000+ annotated instances · roughly six per image. Class presence is indexed frame by frame: class-enriched subsets can be assembled on request. Collection continues under the same six-view protocol and grows every week.

MouthCare technology

Four deep-learning models, trained on this dataset

Every photo is read by four models voting pixel by pixel: only the areas where they converge are flagged. Tartar, cavity, inflamed gum — each area comes back with a probability score, in a visual language the patient understands immediately.

  • Tooth
  • Gum
  • Tartar
  • Cavity
  • Gingivitis
Original intraoral photograph
MouthCare AI segmentation
PATIENT PHOTOREVEALED BY AI
Let's talk

Evaluate the dataset

The sample published on Hugging Face shows exactly the delivered format: 480 images across 80 patients, with annotations, masks and COCO export. The full corpus follows, under agreement.

Dr Romain Lateur · MouthCare — NT Corporation · contact@mouthcare.fr · +33 6 35 16 19 23

Access to the full dataset is agreement-based. Sample images are downloadable only after the evaluation terms are accepted, and no image is sent by e-mail.