- Learning Microsoft Cognitive Services
- Leif Larsen
- 229字
- 2021-08-13 15:40:14
Speech
Adding one of the Speech APIs allows your application to hear and speak to your users. The APIs can filter noise and identify speakers. Based on the recognized intent, they can drive further actions in your application.
The speech domain contains three APIs that are outlined in the following sections.
Bing Speech
Adding the Bing Speech API to your application allows you to convert speech to text and vice versa. You can convert spoken audio to text either by utilizing a microphone or other sources in real time or by converting audio from files. The API also offers speech intent recognition, which is trained by the Language Understanding Intelligent Service (LUIS) to understand the intent.
Speaker recognition
The speaker recognition API gives your application the ability to know who is talking. By using this API, you can verify that the person that is speaking is who they claim to be. You can also determine who an unknown speaker is based on a group of selected speakers.
Translator speech API
The translator speech API is a cloud-based automatic translation service for spoken audio. Using this API, you can add end-to-end translation across web apps, mobile apps, and desktop applications. Depending on your use cases, it can provide you with partial translations, full translations, and transcripts of the translations cover all speech-related APIs in Chapter 5, Speak with Your Application.
- 零點起飛學Xilinx FPG
- Mastering Delphi Programming:A Complete Reference Guide
- 精選單片機設計與制作30例(第2版)
- 筆記本電腦維修300問
- Istio服務網格技術解析與實踐
- Blender Game Engine:Beginner's Guide
- Wireframing Essentials
- Python Machine Learning Blueprints
- The Artificial Intelligence Infrastructure Workshop
- FPGA實驗實訓教程
- DevOps實戰:VMware管理員運維方法、工具及最佳實踐
- PIC系列單片機的流碼編程
- Mastering Unity 2D Game Development
- Hands-On Game Development with WebAssembly
- 電腦組裝與維修實戰