本頁面由 Cloud Translation API 翻譯而成。

Chirp：通用語音模型

Chirp 是 Google 新一代的語音轉文字模型，經過多年研究，Chirp 的第一個版本現已支援語音轉文字功能。我們打算改善 Chirp，並擴大支援更多語言和領域。詳情請參閱「Google USM」一文。

我們訓練 Chirp 模型時，使用的架構與目前的語音模型不同。單一模型可整合多種語言的資料。不過，使用者仍須指定模型應辨識的語言。Chirp 不支援其他模型提供的部分 Google 語音功能。如需完整清單，請參閱「功能支援與限制」。

型號 ID

Chirp 適用於 Speech-to-Text API V2。使用方式與其他模型相同。

Chirp 的模型 ID 為：chirp。

您可以在同步或批次辨識要求中指定這個模型。

可用的 API 方法

Chirp 處理語音時，會將語音切分成比其他模型更大的區塊，因此可能不適合用於真正的即時用途。您可以使用下列 API 方法存取 Chirp：

v2 Speech.Recognize(適合短音訊，長度不到 1 分鐘)
v2 Speech.BatchRecognize(適合 1 分鐘到 8 小時的長篇音訊)

下列 API 方法不支援 Chirp：

v2 Speech.StreamingRecognize
v1 Speech.StreamingRecognize
v1 Speech.Recognize
v1 Speech.LongRunningRecognize
v1p1beta1 Speech.StreamingRecognize
v1p1beta1 Speech.Recognize
v1p1beta1 Speech.LongRunningRecognize

區域

Chirp 適用於下列區域：

us-central1
europe-west4
asia-southeast1

詳情請參閱語言頁面。

語言

如要查看支援的語言，請參閱完整語言清單。

功能支援與限制

Chirp 不支援部分 STT API 功能：

信心分數：API 會傳回值，但並非真正的信心分數。
語音調整：不支援任何調整功能。
說話者區分：系統不支援自動區分說話者。
強制正規化：不支援。
個別字詞信心值：不支援。
語言偵測：不支援。

Chirp 支援下列功能：

自動加上標點符號：標點符號由模型預測。這個機制可停用。
單字時間碼：可選擇是否傳回。
不限語言的音訊轉錄：模型會自動推斷音訊檔案中使用的語言，並將其加入結果。

事前準備

Sign in to your Google Cloud account. If you're new to Google Cloud, create an account to evaluate how our products perform in real-world scenarios. New customers also get $300 in free credits to run, test, and deploy workloads.

In the Google Cloud console, on the project selector page, select or create a Google Cloud project.

Go to project selector

Make sure that billing is enabled for your Google Cloud project.

Enable the Speech-to-Text APIs.

Enable the APIs

Make sure that you have the following role or roles on the project: Cloud Speech Administrator

Check for the roles

In the Google Cloud console, go to the IAM page.
Go to IAM
Select the project.
In the Principal column, find all rows that identify you or a group that you're included in. To learn which groups you're included in, contact your administrator.
For all rows that specify or include you, check the Role column to see whether the list of roles includes the required roles.

Grant the roles

In the Google Cloud console, go to the IAM page.
前往「IAM」頁面
選取專案。
按一下「授予存取權」。
在「New principals」(新增主體) 欄位中，輸入您的使用者 ID。這通常是 Google 帳戶的電子郵件地址。
在「Select a role」(選取角色) 清單中，選取角色。
如要授予其他角色，請按一下「新增其他角色」，然後新增每個其他角色。
按一下 [Save]。

Install the Google Cloud CLI.

Note: If you installed the gcloud CLI previously, make sure you have the latest version by running gcloud components update.

If you're using an external identity provider (IdP), you must first sign in to the gcloud CLI with your federated identity.

To initialize the gcloud CLI, run the following command:

gcloud init

In the Google Cloud console, on the project selector page, select or create a Google Cloud project.

Go to project selector

Make sure that billing is enabled for your Google Cloud project.

Enable the Speech-to-Text APIs.

Enable the APIs

Make sure that you have the following role or roles on the project: Cloud Speech Administrator

Check for the roles

In the Google Cloud console, go to the IAM page.
Go to IAM
Select the project.
In the Principal column, find all rows that identify you or a group that you're included in. To learn which groups you're included in, contact your administrator.
For all rows that specify or include you, check the Role column to see whether the list of roles includes the required roles.

Grant the roles

In the Google Cloud console, go to the IAM page.
前往「IAM」頁面
選取專案。
按一下「授予存取權」。
在「New principals」(新增主體) 欄位中，輸入您的使用者 ID。這通常是 Google 帳戶的電子郵件地址。
在「Select a role」(選取角色) 清單中，選取角色。
如要授予其他角色，請按一下「新增其他角色」，然後新增每個其他角色。
按一下 [Save]。

Install the Google Cloud CLI.

Note: If you installed the gcloud CLI previously, make sure you have the latest version by running gcloud components update.

If you're using an external identity provider (IdP), you must first sign in to the gcloud CLI with your federated identity.

To initialize the gcloud CLI, run the following command:

gcloud init

用戶端程式庫可以使用應用程式預設憑證，輕鬆向 Google API 進行驗證，然後傳送要求給這些 API。使用應用程式預設憑證，您可以在本機測試及部署應用程式，不必變更基礎程式碼。詳情請參閱「驗證以使用用戶端程式庫」。

If you're using a local shell, then create local authentication credentials for your user account:
```
gcloud auth application-default login
```
You don't need to do this if you're using Cloud Shell.

If an authentication error is returned, and you are using an external identity provider (IdP), confirm that you have signed in to the gcloud CLI with your federated identity.

此外，請確認您已安裝用戶端程式庫。

使用 Chirp 執行同步語音辨識

以下是對本機音訊檔案執行同步語音辨識的範例 (使用 Chirp)：

Python

import os

from google.api_core.client_options import ClientOptions
from google.cloud.speech_v2 import SpeechClient
from google.cloud.speech_v2.types import cloud_speech

PROJECT_ID = os.getenv("GOOGLE_CLOUD_PROJECT")


def transcribe_chirp(
    audio_file: str,
) -> cloud_speech.RecognizeResponse:
    """Transcribes an audio file using the Chirp model of Google Cloud Speech-to-Text API.
    Args:
        audio_file (str): Path to the local audio file to be transcribed.
            Example: "resources/audio.wav"
    Returns:
        cloud_speech.RecognizeResponse: The response from the Speech-to-Text API containing
        the transcription results.

    """
    # Instantiates a client
    client = SpeechClient(
        client_options=ClientOptions(
            api_endpoint="us-central1-speech.googleapis.com",
        )
    )

    # Reads a file as bytes
    with open(audio_file, "rb") as f:
        audio_content = f.read()

    config = cloud_speech.RecognitionConfig(
        auto_decoding_config=cloud_speech.AutoDetectDecodingConfig(),
        language_codes=["en-US"],
        model="chirp",
    )

    request = cloud_speech.RecognizeRequest(
        recognizer=f"projects/{PROJECT_ID}/locations/us-central1/recognizers/_",
        config=config,
        content=audio_content,
    )

    # Transcribes the audio into text
    response = client.recognize(request=request)

    for result in response.results:
        print(f"Transcript: {result.alternatives[0].transcript}")

    return response

提出要求並啟用不限語言的轉錄功能

下列程式碼範例示範如何發出要求，並啟用不限語言的轉錄功能。

Python

import os

from google.api_core.client_options import ClientOptions
from google.cloud.speech_v2 import SpeechClient
from google.cloud.speech_v2.types import cloud_speech

PROJECT_ID = os.getenv("GOOGLE_CLOUD_PROJECT")


def transcribe_chirp_auto_detect_language(
    audio_file: str,
    region: str = "us-central1",
) -> cloud_speech.RecognizeResponse:
    """Transcribe an audio file and auto-detect spoken language using Chirp.
    Please see https://cloud.google.com/speech-to-text/v2/docs/encoding for more
    information on which audio encodings are supported.
    Args:
        audio_file (str): Path to the local audio file to be transcribed.
        region (str): The region for the API endpoint.
    Returns:
        cloud_speech.RecognizeResponse: The response containing the transcription results.
    """
    # Instantiates a client
    client = SpeechClient(
        client_options=ClientOptions(
            api_endpoint=f"{region}-speech.googleapis.com",
        )
    )

    # Reads a file as bytes
    with open(audio_file, "rb") as f:
        audio_content = f.read()

    config = cloud_speech.RecognitionConfig(
        auto_decoding_config=cloud_speech.AutoDetectDecodingConfig(),
        language_codes=["auto"],  # Set language code to auto to detect language.
        model="chirp",
    )

    request = cloud_speech.RecognizeRequest(
        recognizer=f"projects/{PROJECT_ID}/locations/{region}/recognizers/_",
        config=config,
        content=audio_content,
    )

    # Transcribes the audio into text
    response = client.recognize(request=request)

    for result in response.results:
        print(f"Transcript: {result.alternatives[0].transcript}")
        print(f"Detected Language: {result.language_code}")

    return response

在 Google Cloud 控制台中開始使用 Chirp

確認您已註冊 Google Cloud 帳戶並建立專案。
前往 Google Cloud 控制台的「Speech」頁面。
如果尚未啟用，請啟用 API。
前往「轉錄稿」子頁面。
按一下「新增語音轉錄作業」
請確認您有 STT 工作區。如果沒有，請建立一個。
1. 開啟「工作區」下拉式選單，然後按一下「新工作區」。
2. 在「建立新工作區」導覽側欄中，按一下「瀏覽」。
3. 按一下即可建立值區。
4. 輸入值區名稱，然後按一下「繼續」。
5. 點選「建立」。
6. 建立值區後，按一下「選取」選取值區。
7. 按一下「建立」，完成建立語音轉文字工作區。
轉錄音訊。
1. 在「新增轉錄稿」頁面中，選擇選取音訊檔案的方式：
  - 按一下「本機上傳」即可上傳。
  - 按一下「雲端儲存空間」，指定現有的 Cloud Storage 檔案。
注意： Speech-to-Text 會嘗試自動評估音訊檔案參數。
1. 按一下「繼續」。
1. 在「轉錄選項」部分，從先前建立的辨識器中，選取您打算用來透過 Chirp 辨識的「說話語言」。
2. 在「模型」* 下拉式選單中，選取「Chirp」。
3. 在「Region」(區域) 下拉式選單中，選取區域，例如「us-central1」。
4. 按一下「繼續」。
5. 如要使用 Chirp 執行第一個辨識要求，請在主要部分中按一下「提交」。
查看 Chirp 轉錄結果。
1. 在「轉錄稿」頁面中，按一下轉錄稿名稱。
2. 在「轉錄詳細資料」頁面中查看轉錄結果，並視需要透過瀏覽器播放音訊。

清除所用資源

如要避免系統向您的 Google Cloud 帳戶收取本頁所用資源的費用，請按照下列步驟操作。

Optional: Revoke the authentication credentials that you created, and delete the local credential file.
```
gcloud auth application-default revoke
```
Optional: Revoke credentials from the gcloud CLI.
```
gcloud auth revoke
```

控制台

In the Google Cloud console, go to the Manage resources page.

Go to Manage resources

In the project list, select the project that you want to delete, and then click Delete.

In the dialog, type the project ID, and then click Shut down to delete the project.

gcloud

Delete a Google Cloud project:

gcloud projects delete PROJECT_ID

後續步驟

練習轉錄短音訊檔案。
瞭解如何轉錄串流音訊。
瞭解如何轉錄長音訊檔案。
如要獲得最佳效能、準確率與其他提示，請參閱最佳做法說明文件。

Chirp：通用語音模型 透過集合功能整理內容 你可以依據偏好儲存及分類內容。

型號 ID

可用的 API 方法

區域

語言

功能支援與限制

事前準備

Check for the roles

Grant the roles

Check for the roles

Grant the roles

使用 Chirp 執行同步語音辨識

Python

提出要求並啟用不限語言的轉錄功能

Python

在 Google Cloud 控制台中開始使用 Chirp

清除所用資源

控制台

gcloud

後續步驟

Chirp：通用語音模型