Case study

A Globally Diverse Selfie Dataset for Face Deduplication

Back to Case Studies

LXT delivered high-volume, demographically balanced selfie dataset to power face deduplication, face recognition and age-estimation models for a leading proof-of-personhood project. 

~500K

images delivered*

20–25K

participants

5

world regions

24

valid images per participant 

6mo

collection window 

Daily

secure delivery via AWS S3 

*Up to 24 valid images per participant across the contracted volume (4 front-facing neutral + 4 head-pose + 16 historical).

The Challenge

An identity system designed to work for every person on earth cannot afford the bias and brittleness that come with web-scraped or studio-style training data. The customer needed a high-volume selfie dataset captured with explicit consent, with enough real-world variation to power face deduplication and identity resolution, face recognition under unconstrained conditions, and age estimation and progression. 

Specifically, the program had to: 

  • Source tens of thousands of consenting adult participants across five world regions 
  • Hit demographic targets on region, age and gender — not just globally, but in a balanced way across the dataset 
  • Capture genuine within-person variation: front-facing neutral captures, structured head-pose captures along two pose axes, and historical imagery showing natural aging over time 
  • Run end-to-end with explicit, linked consent and privacy controls, including full provenance per image 
  • Deliver at the pace of an active model-training program rather than a one-shot batch hand-off 

Off-the-shelf datasets failed the diversity and consent bar; in-house collection at this scale would have absorbed engineering capacity the customer needed to spend on the model itself. 

The Solution

LXT ran the project as a managed, end-to-end image data collection engagement; from purchase order through to acceptance of the final delivery; combining the reach of its global contributor crowd with the controls of an enterprise data operation built for biometric data.

Capturing real within-person variation
A robust facial model needs to see how a single person actually looks across days, devices, lighting, expressions and years — not just a clean studio frame. To get that signal, each participant contributed images in three selfie categories: front-facing captures as clean identity anchors, head-pose captures along an assigned axis to spread the dataset across the pose space, and historical gallery photos that showed natural aging and real-world variation over time. Each participant was assigned exactly one of the two pose variations, and LXT balanced the volume roughly 50/50 across the two — with a small tolerance — so the dataset stayed even at the pose-variation level.

Demographic balance for a globally diverse selfie dataset
Diversity was engineered into the pipeline rather than hoped for after the fact. Quotas were enforced at the regional, age and gender level, and intake was monitored live so the crowd could be steered toward under-represented segments while the program was running — instead of being rebalanced retroactively at the end.

Multi-pass quality assurance
Every submission ran through automated pre-checks first — face detection, image quality, AI/spoof detection, biometric similarity to flag duplicate identities, and metadata consistency — and then through manual review by LXT’s specialist QA reviewers. Structured review and dispute windows gave the customer a clear path to resolve edge cases before billing.

Consent, privacy and security for biometric selfie data
Every participant was an adult who gave clear, informed consent linked to the customer’s privacy agreement before any data was collected. LXT ran the program under ISO 27001-certified processes and GDPR-aligned data handling. Cross-vendor identity deduplication and sub-supplier fingerprinting kept the customer from receiving the same person twice across collection partners, and the data was delivered into a dedicated cloud environment that only the customer could access.

Project Specifications

The following specifications show how LXT structured the selfie dataset for scale, diversity, quality and secure delivery.

  • Volume:
    20,000–25,000 participants · Target of 20 valid images each · Billing cap of 24 per participant · 500K images total
  • Image mix per participant:
    4 front-facing neutral + 4 head-pose + up to 16 historical (minimum 9 valid images for a submission to count)
  • Pose variations:
    Variation A, straight axis: pitch up, left yaw, pitch down, right yaw

    Variation B, diagonal: up-left, up-right, down-left, down-right · One variation per participant, balanced 50-50 (with 5% tolerance) across the dataset.
Selfie Dataset example
Selfie dataset example (Variation B)
  • Historical time span:
    Min. 12 months and max. 20 years between the oldest historical image and the newest front-facing capture; no more than one historical image per calendar day
  • Regional targets:
    ~20% each from Europe/North America/Australia, Africa, South Asia, East/Southeast Asia, South America (min. 10% per region)
  • Age and gender targets:
    Five age brackets (18-24, 25-34, 35-44, 45-54, 55+) bounded between 5% and 30% per bracket · Gender balance held above a 45% floor on the under-represented gender
  • Eligibility & consent:
    18+ at time of collection, with explicit informed consent linked to the customer’s privacy agreement before any image is captured
  • Capture interface:
    Front camera at arm’s length for current captures via the LXT contributor app (smartphone) or workplace (desktop). Historical photos can be uploaded from the participant’s gallery.
  • Delivery:
    Daily upload into a dedicated cloud environment with exclusive customer access
  • Engagement length:
    Six months from PO to acceptance of final delivery

Contributor Task Flow

Each participant moves through a structure task on the LXT platform:

  1. Self-declare eligibility and demographic information (region, age bracket, gender)
  2. Confirm and give explicit consent to the customer’s privacy agreement and biometric data terms
  3. Capture of front-facing neutral photos with the front camera at arm’s length, varying clothing, lighting and background 
  4. Capture of head-pose photos along the assigned pose variation (Variation A — straight axis, or Variation B — diagonal axis) 
  5. Upload a gallery of historical photos covering at least one year and up to twenty years of natural variation over time
  6. Submission for QA review, with a structured dispute path for any rejections 

Outcomes

The engagement produces the artifacts a modern facial-AI program actually needs:

  • A globally diverse, consent-backed selfie dataset built to explicit regional, age and gender targets — with documented provenance for every image 
  • Three selfie image categories per participant: front-facing, head-pose, historical delivered consistently per participant 
  • Daily secure delivery into the customer’s AWS environment 
  • Explicit linked consent, ISO 27001-certified processes, GDPR-aligned handling, and cross-vendor identity deduplication 

Why LXT

When a system has to work for everyone, who collects the data matters as much as the data itself. The customer chose LXT for:

  • True global reach — a crowd that genuinely spans 150+ countries and 1,000+ language locales, not just a handful of major markets 
  • A repeatable diversity playbook for facial imagery, with quotas enforced at the regional, age and gender level 
  • Enterprise-grade QA: automated pre-checks plus specialist manual review, with structured dispute handling 
  • ISO 27001-certified processes, GDPR-aligned consent capture, and cross-vendor identity deduplication built into the workflow 
  • Operational maturity to run a multi-region program at the pace of an active training pipeline