OneRuby.devAN ENGINEERING NOTEBOOK

ruby · 5 min read

Speech Recognition in Ruby: Send Audio Bytes, Handle Every Result

Build a Google Speech-to-Text V1 request from local audio bytes and extract segmented results, with real Ruby protobuf and SDK method tests offline.

A local filename is not an audio upload. If a request puts /tmp/meeting.wav in Google Speech-to-Text's audio.uri, the service cannot read that file from your computer. The field expects a Cloud Storage URI. For a local clip, the Ruby client should send the file's bytes through audio.content.

That boundary is easy to test without paying for recognition. This example uses the actual Google Speech-to-Text V1 Ruby types and recognition method, replacing only the transport with a synthetic response fixture. It checks request construction and transcript extraction. It does not measure speech accuracy or establish that an account can call the hosted service.

Pick an API and an audio contract

The attachment pins google-cloud-speech-v1 1.8.1. These are V1 messages and methods; a V2 recognizer resource and request have a different shape. Keep the version visible when consulting documentation or copying another example.

For the small synchronous path here, the input contract is a header-bearing WAV or FLAC file. The V1 configuration reference allows encoding and sample rate to come from those headers. Renaming an MP3 to .wav does not convert it, and setting a sample-rate field does not resample its samples.

The request builder reads binary data rather than text:

Ruby
def self.request_for(path, language: 'en-US')
bytes = File.binread(path)
raise ArgumentError, 'empty audio' if bytes.empty?
# This example accepts header-bearing WAV/FLAC; let the service read their header.
Google::Cloud::Speech::V1::RecognizeRequest.new(
config: {language_code: language, enable_automatic_punctuation: true},
audio: {content: bytes})
end

This helper checks that the file exists and has bytes. It does not validate its container, duration, codec or channel layout. A real upload path should inspect those properties before choosing recognition settings and enforce its own size limits before reading the entire file into memory.

Ruby protobuf messages take raw bytes for content. Manually base64-encoding those bytes would change what the SDK receives. Base64 is relevant to the REST JSON representation; the RecognitionAudio reference also describes the mutually exclusive content and URI choices. The fixture round-trips the Ruby request through protobuf encoding to check that the audio bytes survive unchanged.

Results are segments; alternatives are candidates

A response can contain several results, each with competing transcript alternatives. Taking only the first result loses later segments. Joining every alternative invents a transcript containing multiple guesses for the same segment.

The extraction rule here selects the first alternative from each result and joins those segments:

Ruby
response.results.filter_map do |result|
alternative = result.alternatives.first
alternative&.transcript unless alternative&.transcript.to_s.empty?
end.join(' ')

The fixture includes “Hello, Ruby.” and its competing “Hello, Rudy.” alternative, followed by “Next sentence.” The returned text is “Hello, Ruby. Next sentence.” No confidence threshold is applied. Confidence values are not a substitute for an evaluation set, and this test says nothing about which hypothesis the recognizer would produce for a real recording.

An empty response or a result with no alternatives produces an empty string instead of a nil dereference. Whether that becomes “no speech detected”, a retry prompt or a review queue is an application decision. Do not silently turn authentication failures, timeouts or invalid-audio errors into that same empty-success state.

The recognize response reference describes the ordered results and alternatives that this extraction follows. More complex jobs may need channel tags, word timestamps and speaker information retained as structured data instead of flattening everything into one string.

What the offline test actually runs

Download the example and locked dependencies. After installing the bundle, run bundle exec ruby test_speech.rb.

Seven tests passed on Ruby 3.3.2. The fixture creates a valid 100 ms mono PCM WAV containing silence. That is a deterministic byte fixture, not a recording from which a service produced the canned words. Tests cover protobuf round-tripping, content-versus-URI behavior, multiple results, competing alternatives, empty results, an empty file and a missing file.

The test invokes the real SDK's recognize method with a configured client object, then records the call at its transport boundary. It deliberately bypasses credential discovery and normal client initialization. This is a narrow, version-pinned adapter test; it does not test gRPC networking, authentication, API quotas or service-side audio validation. A dependency update that changes these internals requires revisiting the adapter.

For an actual request, construct Google::Cloud::Speech::V1::Speech::Client.new using the documented authentication setup and pass it to SpeechClip.transcribe(client, path). That live path is intentionally separate from the test suite. The Ruby client reference covers client construction and recognition calls.

The next test needs real speech

Before using this in an application, choose short recordings whose transcripts you know and have permission to process. Include silence, background noise, names and the languages your users actually speak. Record the requested configuration alongside each result so a later change can be compared meaningfully.

Check the current service limits when deciding between synchronous recognition and an asynchronous or streaming path. Long recordings also need a job lifecycle: stable identifiers, progress, retries and retention decisions. An offline green test cannot answer those questions.

The useful separation is now clear: bytes and response handling have deterministic checks; recognition quality and hosted-service access still require an explicit live experiment. Keeping those claims separate makes a working upload easier to debug when the words are wrong.

Found a mistake or tried a different approach?

Send Alex a note ↗