Practice Exams:

Microsoft AI-103: Speech-to-Text, Text-to-Speech, and Translation

AI & Machine Learning

Speech interfaces turn natural conversation into application input, but reliability depends on what happens before and after the model produces words. Microsoft AI-103 includes speech-to-text, text-to-speech, speech translation, audio as an agent modality and custom speech capabilities where supported. A production developer must choose a service for each job, measure recognition and synthesis quality, manage data and consent, and keep agents from acting on uncertain transcripts as though they were verified commands.

On this page
  1. Choose the speech interaction the application needs
  2. Build speech-to-text with useful context
  3. Treat speakers and conversation boundaries carefully
  4. Generate speech and evaluate synthesized voice
  5. Translate text and speech while preserving meaning
  6. Connect audio to agent tools without unsafe actions
  7. Monitor quality, cost and privacy
  8. Prepare exam scenarios using explicit evidence

Choose the speech interaction the application needs

A call-center transcript, a voice assistant and a translated meeting require different designs. Batch transcription prioritizes throughput, timestamps and speaker context; live dictation prioritizes latency and partial results; a spoken agent needs both recognition and response generation, often under interruption. Specify languages, accents, background noise, expected vocabulary, concurrency and acceptable delay before selecting an SDK or architecture. In the AI-103 blueprint, speech is part of implementing text analysis and agentic modalities rather than an isolated telephony specialty.

Check the current supported options in Microsoft’s Azure Speech documentation. Availability of language, voice, customization, transcription method and region changes over time; do not assume the same service accepts every audio file or supports the same real-time streaming mode. A supported feature in one language may not be available in another.

Before implementing a lab, write down the chosen language locale, selected Azure Speech resource region, audio format, duration and whether the task needs word-level timing, diarization or low latency. Confirm each capability in the current service documentation rather than assuming a feature in batch transcription exists in every streaming path. A five-minute recorded interview should not be processed with a single-utterance demonstration API designed for brief speech.

Build speech-to-text with useful context

A recognizer receives audio but application developers determine sampling, encoding, channel handling, segmentation and postprocessing. If a transcript will be used to search documents or authorize an action, numbers, names, acronyms and negation deserve special attention. A system that turns “do not renew” into “renew” has changed the user’s intent. Preserve a traceable link between transcript spans and original audio, and allow correction before a high-impact step.

For a lab, transcribe samples with different microphones, noise levels, accents and technical terms. Compare word error rate or task-specific critical-word accuracy with human reference transcripts. Also measure time to first partial result, final-result delay and how often speaker turns are incorrectly combined. A fluent transcript can still be wrong. Separate the speech component’s accuracy from the downstream language model’s ability to respond appropriately.

Speech input should have an explicit end-of-utterance strategy. In a quiet dictation workflow a short pause may be enough, but in a conversation with slow speakers the same timeout can truncate a sentence. Test silence thresholds and interruption handling against representative recordings. If the recognizer emits partial results, avoid pushing every unstable fragment into a tool call. Wait for a confirmed result or a user review before using critical values as business-action parameters.

For a distributed application, plan how to reconnect after a stream drops. The client should tell the user what was received, whether a final transcript was saved, and whether further processing occurred. Deduplicate segments when a connection resumes so one sentence is not charged or acted on twice. Monitor audio duration separately from transcript length, because quiet intervals, background noise and retransmission can affect cost and latency differently.

Microsoft’s Speech-to-text Python quickstart uses the azure-cognitiveservices-speech package. For a short local WAV file, configure SPEECH_KEY and ENDPOINT as environment variables in a controlled training environment; do not commit their values to source control. The sample below uses the documented single-utterance operation and exposes non-success outcomes instead of printing an empty transcript as if recognition passed:

import os
import azure.cognitiveservices.speech as speechsdk

cfg = speechsdk.SpeechConfig(
    subscription=os.environ["SPEECH_KEY"],
    endpoint=os.environ["ENDPOINT"],
)
cfg.speech_recognition_language = "en-US"
audio = speechsdk.audio.AudioConfig(filename="sample.wav")
recognizer = speechsdk.SpeechRecognizer(
    speech_config=cfg, audio_config=audio,
)
result = recognizer.recognize_once_async().get()
if result.reason == speechsdk.ResultReason.RecognizedSpeech:
    print(result.text)
elif result.reason == speechsdk.ResultReason.NoMatch:
    print("No recognized speech")
elif result.reason == speechsdk.ResultReason.Canceled:
    details = result.cancellation_details
    print("Recognition canceled:", details.reason)
    if details.reason == speechsdk.CancellationReason.Error:
        print("Service error:", details.error_details)
else:
    print("Unexpected recognition result:", result.reason)

This quickstart is intentionally limited to an utterance of roughly 30 seconds or until silence ends it. For long recordings use the documented continuous, fast or batch alternatives. The exercise should include a clear sample, a silence-only file and an unsupported or incorrectly encoded file; record the result category and compare recognized words with a verified transcript. Use keyless identity where supported in production rather than treating a training key pattern as a production default.

Treat speakers and conversation boundaries carefully

Multi-speaker meetings create ambiguity about who said what. Where diarization is supported, keep speaker labels as an inference that may need review rather than asserting legal identity from voice characteristics. If an application summarizes decisions or produces action items, preserve relevant timestamps and provide a way to open the source audio. A single meeting can contain a proposal and a later rejection; an inaccurate summary can turn the proposal into an approved instruction.

Voice agents should preserve explicit conversation state: partial user utterances, interruptions, turn boundaries and confirmations. A user may start to provide an address and then correct it. The final agent action must use the confirmed value, not the first partial transcript. Test barge-in and silence handling as well as ordinary completed turns; otherwise low-latency synthesis may speak confidently over a correction.

Speaker labels are hypotheses, not identity authentication. In a two-speaker meeting test, annotate each turn by listening to the recording and measure both word recognition and speaker attribution. Mark interruptions separately from completed utterances. A correct-looking transcript with the two speakers swapped is unfit for automated attribution of commitments or approvals.

Generate speech and evaluate synthesized voice

Text-to-speech should produce clear, context-appropriate output, including numbers, units, names and safety notices. A realistic voice is not sufficient if it pronounces an address or medication incorrectly. Choose supported voices and output formats according to channel bandwidth and user accessibility needs. Provide text alternatives, captions or transcripts for users who cannot hear the output, and consider whether a slower speaking rate or explicit pronunciation guidance is needed.

If custom voices are used, follow consent, rights and service-specific approval requirements rather than treating any sample voice as available for cloning. Test whether synthesized speech sounds acceptable over real devices, not only high-quality headphones. In an agent, keep the spoken response aligned with the displayed text and the action actually taken. A voice can make a hallucinated answer seem persuasive; synthesis quality and answer truthfulness must be measured separately.

For text-to-speech, include a pronunciation test for an acronym, a date, an amount and a long technical term. Retain the input text, selected locale and voice, and a human judgment of intelligibility; simply receiving an audio file does not prove correct pronunciation. Custom voice work also requires confirming the applicable Microsoft access, consent and rights requirements before collecting speaker recordings.

Translate text and speech while preserving meaning

Speech translation is more than recognizing words in one language and paraphrasing them in another. Important details include negation, speaker context, product names, regional conventions, proper nouns and culturally specific phrases. A general-language model may provide natural phrasing but subtly change legal or technical meaning. For critical workflows, use terminology controls, bilingual review and traceable source segments. Test whether the system preserves amounts, dates and identifiers exactly when instructed to do so.

Choose between supported Translator capabilities, speech-specific translation functions and LLM-powered translation based on language coverage, style needs, auditability and latency. Do not assume every speech voice supports translation to every target language. If the source cannot be recognized reliably, show uncertainty rather than silently translating a guessed transcript.

Connect audio to agent tools without unsafe actions

A spoken assistant may route a recognized request to retrieval or an API. That creates a trust boundary between recognition and authorization: the user’s voice input is data to interpret, not automatic permission to issue a privileged write. Ask for an explicit confirmation before high-impact actions and repeat critical parameters back in an accessible way. Verify the authenticated user through the application’s identity mechanism; do not rely solely on the fact that a voice was recorded.

For a lab, submit a command with a misrecognized name and a noise-induced amount change. The safe agent should query or confirm before acting, then log what value was approved. Evaluate whether the tool schema rejects missing or invalid fields. The trust controls of agent tool calling and authorization must reject low-confidence spoken commands under the same approval rules as typed requests.

Monitor quality, cost and privacy

Speech services handle recordings that may contain private conversations, biometric-like characteristics or regulated personal data. Classify what is collected, decide whether raw audio must be retained and for how long, restrict operator access and avoid logging full recordings by default. Track recognition failures, latency, cost per audio duration, unsuccessful translation requests and actionable user corrections. A low model error rate is not enough when critical names remain wrong.

Test failure paths that operators will actually see: unsupported language, bad file encoding, interrupted stream, expired credentials and an unavailable downstream model. The system should provide a helpful error and avoid creating a partially completed action when input processing fails. Alerts should identify the failing layer—ingestion, recognition, translation, reasoning or speech output—rather than reporting all failures as a generic AI exception.

Prepare exam scenarios using explicit evidence

A useful Microsoft AI-103 lab includes one batch transcript, one streaming or interactive request where supported, a synthesized response, and one translation experiment. Compare outputs with reference text and record the supported language and deployment mode used. For audio-powered agents, add a confirmed-action workflow with a negative authorization test. The exam covers choosing and integrating these capabilities, so an informed tradeoff is more valuable than memorizing every SDK method name.

Recorded-media analysis has a different result contract from interactive speech: with audio and video Content Understanding, recordings become timestamped evidence segments that can be inspected or searched, whereas recognition yields words, translation changes languages, and synthesis returns audio. Applications that combine these services require distinct validation and handoff schemas at each stage.

Continue learning

Related guides

Content Understanding for Messy DocumentsBusiness documents rarely arrive as clean paragraphs ready for a language model.Evaluating Agents for Accuracy and SafetyAgent evaluation is more complicated than checking whether a model produced the expected sentence.Securing Azure AI EndpointsSecuring an Azure AI endpoint requires several controls working together: authentication, authorization, network exposure, rate limits, request validation, monitoring, and deployment…Tool Calling in Azure AI AgentsTool calling turns an Azure AI agent from a conversational model into a system that can retrieve live information, call APIs, run business operations, or interact with external services.

Related Posts

• Mastering the AI-102 Exam: Your Azure AI Engineer Associate Roadmap

• Mastering AI-102: A Complete Preparation Resource

• Understanding the Core of AI-102 and the Azure AI Engineer Role

• Microsoft AB-100: AI Across Dynamics 365, Power Platform, and Foundry

• Microsoft AI-103: Rate Limits, Cost, and Scaling Azure AI Applications

• Microsoft AI-103: Blue-Green Releases for AI Endpoints

• Microsoft AI-103: Online Evaluation for AI Systems

• Microsoft AI-103: Securing Azure AI Endpoints

• Microsoft AB-100: DLP Policies for Copilot Studio

• Microsoft AB-100: Human Handoff in Copilot Studio