Free tools Windows power users keep installed
One-click scans. No signup required.
To use Google Cloud Text-to-Speech in Java, enable the API and billing on a Google Cloud project, authenticate with Application Default Credentials (ADC), then use the google-cloud-texttospeech client to send text or SSML and save the returned audio bytes. The example below creates an MP3 file; you can adapt its voice, encoding, and output handling to your application.
1. Prepare Google Cloud and Java credentials
Start with a Google Cloud project that has billing configured and the Cloud Text-to-Speech API enabled. Google’s client-library quickstart also walks through installing the Google Cloud CLI and running gcloud init.
For a local development shell, configure ADC with:
gcloud auth application-default login
Java client libraries use ADC so the application can obtain credentials from its environment rather than embedding credentials in source code. Use an appropriate production identity and runtime credential source in production; the authentication approach can change between environments without changing the synthesis call. See Google’s Java authentication guidance.
2. Add the Java client library
For Maven, Google’s quickstart shows the Google Cloud libraries BOM and the Text-to-Speech artifact. The versions shown there are examples captured from the documentation, not a guarantee that they are still the newest; check the page when setting up a new project.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →<dependencyManagement>
<dependencies>
<dependency>
<groupId>com.google.cloud</groupId>
<artifactId>libraries-bom</artifactId>
<version>26.86.0</version>
<type>pom</type>
<scope>import</scope>
</dependency>
</dependencies>
</dependencyManagement>
<dependencies>
<dependency>
<groupId>com.google.cloud</groupId>
<artifactId>google-cloud-texttospeech</artifactId>
</dependency>
</dependencies>
Managing Google Cloud libraries through the BOM avoids specifying a separate version for this artifact in the dependency declaration. The quickstart also gives Gradle and sbt examples; its sbt example displays google-cloud-texttospeech version 2.99.0. Check the current quickstart for the syntax and versions appropriate to your build tool.
3. Synthesize text and write an MP3
A synthesis request has three distinct parts: the input text, voice selection, and audio configuration. This minimal example follows the official Java sample and writes the binary response content to output.mp3:
Rank #2
import com.google.cloud.texttospeech.v1.AudioConfig;
import com.google.cloud.texttospeech.v1.AudioEncoding;
import com.google.cloud.texttospeech.v1.SsmlVoiceGender;
import com.google.cloud.texttospeech.v1.SynthesisInput;
import com.google.cloud.texttospeech.v1.SynthesizeSpeechResponse;
import com.google.cloud.texttospeech.v1.TextToSpeechClient;
import com.google.cloud.texttospeech.v1.VoiceSelectionParams;
import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
public class TextToSpeechExample {
public static void main(String[] args) throws IOException {
try (TextToSpeechClient client = TextToSpeechClient.create()) {
SynthesisInput input = SynthesisInput.newBuilder()
.setText("Hello, World!")
.build();
VoiceSelectionParams voice = VoiceSelectionParams.newBuilder()
.setLanguageCode("en-US")
.setSsmlGender(SsmlVoiceGender.NEUTRAL)
.build();
AudioConfig audioConfig = AudioConfig.newBuilder()
.setAudioEncoding(AudioEncoding.MP3)
.build();
SynthesizeSpeechResponse response =
client.synthesizeSpeech(input, voice, audioConfig);
Files.write(Path.of("output.mp3"),
response.getAudioContent().toByteArray());
}
}
}
The client is created with TextToSpeechClient.create(), which uses ADC. The request supplies plain text, an en-US language code and neutral gender hint, and MP3 encoding. The response’s audio content is binary data; converting it to a byte array lets Java write it directly to a file. Keep the output extension consistent with the encoding requested.
4. Choose plain text or SSML
Plain text is the simplest input for straightforward narration. Use SSML when the spoken result needs explicit markup for pronunciation, pauses, emphasis, dates, or addresses. SSML is only the input: voice choice and output encoding remain separate fields in the request.
To submit SSML, build the input with setSsml instead of setText:
String ssml = "<speak>Hello, <break time="500ms"/> world.</speak>";
SynthesisInput input = SynthesisInput.newBuilder()
.setSsml(ssml)
.build();
Ensure the markup is well formed and follows the W3C Speech Synthesis specification. Google’s SSML sample shows the Java input pattern and notes that a voice may be specified by name.
Rank #4
5. Select a voice and audio format
Voice selection
The example uses a language code and gender hint. For more specific voice selection, set a voice name after finding it in Google’s current supported voices and languages catalog. Voice names, language availability, and voice families can change, so verify the catalog rather than relying on a hard-coded assumption.
Audio encoding and storage
The example requests MP3, but the audio configuration is a separate request field and can specify another supported encoding. The AudioConfig reference documents the available configuration. Once returned, the binary audio can be written to a file, stored in object storage, or passed into an application’s media pipeline; choose handling suited to your application and the selected encoding.
Quick Recap
Best Value
6. Troubleshoot common setup problems
- Authentication fails: For a local shell, confirm that
gcloud auth application-default logincompleted and that the intended project is configured. In production, make sure the runtime identity is available through ADC. - The request is rejected before synthesis: Check that the Text-to-Speech API is enabled for the project and billing is configured.
- The voice is unavailable: Verify the language code and voice name against the current supported-voices catalog; do not assume a voice identifier remains available.
- The output cannot be played: Confirm the requested encoding matches how the bytes are saved and the file extension used. The response contains binary audio, not text to print or append to a string.
- The build cannot resolve the library: Verify the artifact coordinates and current dependency-management example for your Maven, Gradle, or sbt configuration.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

