Principal Speech Data Linguist

Innodata Inc. · Remote - United States

  • Senior
  • Full-time
  • $160,000 – $185,000
  • Posted 2026-09-24
  • Confirmed live on 25 September 2026

Apply at Innodata Inc.

Job description

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role:

As speech and audio models get better, the human role gets harder, not easier — it moves from producing transcripts to defining what a correct one is, adjudicating the cases models still get wrong, and designing the human-in-the-loop workflows that keep improving them. Innodata runs high-volume segmentation and transcription workflows for the customers and frontier labs building these models, and we are hiring a principal-level linguist to own the linguistic standards and quality behind that work — today, and as the workflows evolve alongside the models over the next two years.

This is the applied-expert counterpart to our Speech & Audio Research Scientist. You set the standards the models are trained and measured against, and you understand the big picture: how different transcription and segmentation methods change what a model learns, and how that ripples into the speech and content-understanding systems our partners are building. You know the research and you know the tools — from IPA and acoustic analysis to forced alignment and the ASR engines our partners benchmark against — but your leverage is linguistic judgment and standard-setting at scale, not building models yourself.

What You’ll Own:

• You will own the linguistic foundation of Innodata's segmentation and transcription work across languages, domains, and use cases. Concretely, you will:

• Define transcription and segmentation standards, style guides, and annotation conventions — verbatim and clean/intelligent verbatim, IPA and phonetic transcription, timestamping and boundary segmentation, speaker labeling and diarization labels, disfluencies and non-speech events, code-switching, and orthographic conventions.

• Establish and run the quality frameworks behind that work: rubrics, error taxonomies, adjudication processes, inter-annotator agreement, and human QA at scale.

• Own the quality lifecycle for transcription and segmentation deliverables end to end — pre-processing and normalization of incoming data, quality checks at the point of acceptance, post-processing and pre-delivery validation against spec, and report creation and packaging for delivery — partnering with delivery operations on execution at scale.

• Design the human-in-the-loop workflows themselves — deciding where human review, correction, and adjudication add the most value as ASR quality rises, so our experts spend their time on what the models still can't do rather than on what they already can.

• Handle the linguistically hard cases models fail on — accented and dialectal speech, low-resource and multilingual audio, overlapping speech, domain jargon (medical, legal, technical), and noisy acoustic conditions.

• Partner with the Speech & Audio Research Scientist to turn model objectives into transcription and segmentation specifications, and to work out how different transcription methods — verbatim versus clean, phonetic versus orthographic, and how audio is segmented and labeled — affect the training and evaluation of ASR, TTS, and speech and content-understanding models.

• Train, calibrate, and mentor expert transcribers and reviewers, and build the onboarding and calibration that keep quality consistent as the work scales.

• Represent Innodata's transcription and segmentation approach to the customers and frontier labs we partner with, and contribute to the methodology and best-practice documentation that make our work legible to their teams.

You’ll Thrive in This Role If You Have:

• Substantial industry experience (typically 8+ years) in transcription, segmentation, and speech-data quality — enough that you have authored standards, not only followed them. This is a principal-level role, and we weight practical depth heavily.

• A Bachelor's degree in linguistics, phonetics, or computational linguistics, or a closely related field, is required — with a strong foundation in phonetics, phonology, and sociolinguistics so that IPA, prosody, disfluency, dialect, and register are native concepts. An advanced degree is preferred.

• A big-picture grasp of how transcription and segmentation choices flow downstream into modeling — how different methods change what speech and content-understanding models learn, and therefore which method fits which modeling objective. You can explain to a model builder why a transcription decision matters.

• Fluency i

Prepare for the interview

Nothing collected for this employer yet. The Blind 75 is what technical screens draw from; practise it here, with a coach, in Java or Python.

More at Innodata Inc.

All open software jobs