Open roles

· Speech Data, Quality & Annotation

Audio QA Lead

Besimple AI

Senior · 5+ yrsRemote · WorldwidePart-time contract (can turn into full-time)Listed 14d ago
Apply now

Backed by

Y Combinator

HQ

🇺🇸 San Francisco

Open roles

4

Experience Senior · 6+ yrs (5+ years)

About the role

structured by ORI

Back to jobsAbout the roleWe are hiring an Audio QA Lead to support the development of high-quality training datasets for next-generation voice AI models.In this role, you will work hands-on to improve the quality, consistency, and usability of speech datasets across applications such as text-to-speech,…

What you will do

  • Develop, refine, and apply audio quality guidelines for speech and voice datasets.
  • Review audio files against technical, linguistic, and task-specific standards, making clear approval, rejection, or revision decisions.
  • Identify audio and annotation issues such as background noise, clipping, distortion, plosives, echo, low signal, segmentation errors, transcript mismatches, and speaker-label inconsistencies.
  • Perform annotation and QA tasks, including transcription, timestamp validation, VAD/segmentation, diarization, pronunciation checks, and metadata review.
  • Record speech based on provided scripts and performance guidelines, delivering natural, high-quality, specification-compliant audio.

What they are looking for

  • Direct experience working with audio AI training datasets or evaluation workflows.
  • Hands-on experience with TTS, ASR, transcription, speech-to-speech, or related voice AI systems.
  • Experience developing or applying audio quality standards in production environments.
  • Experience with speech annotation tasks such as transcription, timestamp QA, VAD/segmentation, and diarization.
  • Strong auditory judgment with the ability to consistently identify subtle audio quality issues.
  • Ability to produce high-quality recordings in a controlled, quiet environment using professional or near-professional equipment.
  • Strong written communication skills with the ability to provide clear, actionable feedback.
  • High attention to detail and sound judgment when evaluating edge cases.
  • Comfort working with structured data formats such as spreadsheets, CSV, or JSON.

Nice to have

  • Experience with audio tools such as Audacity, Praat, or similar.
  • Basic scripting skills in Python, Bash, or SQL for QA or dataset analysis.
  • Background in linguistics, phonetics, speech research, or voiceover work.
  • Experience evaluating both real and synthetic audio.
  • Multilingual experience or familiarity with accents and dialect variation.
  • Familiarity with compliant handling of consented and licensed voice data.
TTSASRtranscriptionspeech-to-speechVAD/segmentationdiarizationlinguisticsphoneticsCSVJSONAudacityPraatPythonBashSQL
Full posting text

Define and apply audio quality standards, record high-quality speech on demand, and perform annotation and QA across speech datasets.

Apply