Skip to main content

Data Collection Synthetic Voice Generative AI

Comparative Evaluation of Voice AI Models

The Challenge

Our client, a global leader in AI computing, required a highly specialized dataset to evaluate the quality of machine-generated audio. For each task, human evaluators had to assess five audio samples on a 0-100 scale across characteristics ranging from technical audio defects to highly subjective elements such as emotion and likability.

The client provided raw data and brief descriptions of each evaluation category but did not have detailed guidelines, benchmark examples, or an established gold-standard dataset. They needed DataForce to:

  • Turn vague requirements into a structured evaluation framework 

  • Recruit and qualify a specific demographic of US-based evaluators

  • Establish agreement across objective and subjective criteria

  • Automate the pipeline from contributor sourcing through final delivery

• • • •The Solution• • • •

DataForce developed a customized human-in-the-loop evaluation program combining expert-led guideline development, automated contributor qualification, and continuous quality monitoring. The team first established comprehensive guidelines and a benchmark audio library so evaluators had reference points.

An automated pipeline processed 2,127 applicants and identified 458 qualified participants while using only 20 of the 75 recruitment hours allocated to the project. Each candidate completed a multi-stage qualification process that included:

  • A hearing assessment requiring recognition of at least 8 out of 10 tones
  • A guideline quiz requiring a minimum score of 85%
  • A training task using representative audio samples
  • An inter-annotator agreement assessment requiring a minimum score of 75%

To manage the levels of subjectivity across audio characteristics, the team developed a subjectivity matrix that established appropriate agreement expectations for each category. Objective characteristics, such as noise and hissing, required high agreement, while subjective metrics, including likability and emotion, allowed for greater variation within controlled limits.

The team monitored data reliability through:

  • Krippendorff’s alpha
  • Fleiss’ kappa
  • ICC2k
  • Reference agreement of at least 85%
  • Overall evaluator agreement of at least 75%

The DataForce Platform automated the workflow from qualification through production and final delivery. Built-in quality controls monitored scoring patterns, excessive use of extreme ratings, suspicious task-completion times, and IP address inconsistencies. Contributors received immediate feedback during training, while those who did not meet the required thresholds were automatically prevented from advancing to production.

By reducing manual file handling and administrative coordination, the workflow protected data integrity, minimized operational risk, and enabled DataForce to deliver consistent, high-quality audio evaluations at scale.

The Results

DataForce delivered the full dataset of 350 audio evaluation tasks ahead of schedule, with zero client rejections and no rework required. The client received a complete, quality-validated dataset without delays, gaining detailed insights into where its generative audio models performed well and where further refinement was needed. 

By combining structured guidelines, carefully qualified human evaluators, and automated quality controls, DataForce gave the client a reliable framework for comparing model performance and identifying targeted opportunities for improvement. 

Appendix A: Comparative Model Performance Dashboard

 

Appendix B: Inter-Annotator Agreement & Calibration Matrix

 

Appendix C: Statistical Reliability & Category Analysis

Request a consultation.