ruby · 5 min read
Amazon Transcribe in Ruby: Separate Starting, Waiting and Reading
Separate Amazon Transcribe submission, bounded polling and authenticated S3 output reading in Ruby, with tested SDK stubs for failures and empty results.
A transcription job has at least three different failure boundaries: submitting the job, waiting for completion, and reading the output. Retrying the whole workflow whenever any one step fails can create duplicate jobs or hide a permissions problem as an empty transcript.
This Ruby example gives those operations separate methods. The tests use the real AWS SDK with its response-stubbing mode, so requests and response types are exercised without uploading audio or calling AWS. The result is a checked orchestration example, not a live transcription or IAM test.
Give a job a durable identity
Amazon Transcribe batch jobs read audio from S3. The request needs a job name and language choice, and this example selects an output bucket and an exact JSON key. The caller supplies those values:
def start(name:, input_uri:, output_bucket:, output_key:) @transcribe.start_transcription_job( transcription_job_name: name, language_code: 'en-US', media: {media_file_uri: input_uri}, output_bucket_name: output_bucket, output_key: output_key) endGenerate a unique name when an application creates a transcription record, then persist it with the input object and requested settings. Do not generate another name each time a worker resumes. A timestamp rounded to seconds is not a reliable unique identifier for concurrent work.
A duplicate-name conflict is not a universal idempotency guarantee. If submission times out ambiguously, look up that persisted job name and verify that the existing job belongs to the intended operation. A different input under the same name must not become an accepted success. This sample surfaces SDK errors; it does not implement that reconciliation policy.
The StartTranscriptionJob reference documents the job fields and output-key behavior. The example fixes the language to en-US and lets the service inspect media properties. It does not demonstrate language identification, channel separation or speaker diarization.
Poll the existing job, with an exit
The waiting method receives the same name. Its loop has four possible outcomes:
case job.transcription_job_status when 'COMPLETED' then return job when 'FAILED' then raise Failed, job.failure_reason.to_s when 'QUEUED', 'IN_PROGRESS' @sleeper.call(interval) if index + 1 < attempts else raise Failed, "unexpected status #{job.transcription_job_status.inspect}" endA completed response returns the job. A failed response preserves its reason. A queued or running job consumes one poll, and an unfamiliar state raises an error rather than spinning forever. After the configured number of attempts, StillRunning carries the job name so a scheduler can check it again later.
An attempt limit bounds the number of calls, not exact elapsed wall time. Network timeouts and SDK retries can extend a live call. Set those client options and an application deadline according to the worker's budget. This short example is easier to inspect with a count limit than with an invisible endless loop.
A throttled poll must not call start again. The test injects a LimitExceededException and verifies that only get_transcription_job was requested. In a real worker, retry the poll with an appropriate delay or reschedule it using the persisted name. Keep the retry decision at the operation that failed.
Read your private output with the S3 client
With a customer-owned output bucket, knowing a URL does not grant permission to read an object. This example reads the explicit bucket and key through Aws::S3::Client#get_object, using the client's authentication rather than an unauthenticated HTTP fetch.
The output parser extracts results.transcripts[*].transcript. It joins complete transcript entries and does not rebuild punctuation from individual word items. An empty transcript list returns an empty string. Malformed JSON, missing required keys and an S3 AccessDenied error remain failures.
Amazon also supports service-managed output with a temporary download URI. That is a different retrieval path; the Transcript reference explains the distinction. Do not take a downloader written for a temporary service-managed URI and assume it can read every private S3 result.
The wrapper reads the JSON body into memory. An application should bound expected output size and decide where to retain raw results, redacted results and audio. Those choices are part of the product's data handling, not properties established by this parser.
Exercise the states without a cloud account
Download the example, tests and locked gems, install the bundle, and run bundle exec ruby test_transcription.rb.
The test clients set stub_responses: true. AWS documents that this disables network traffic and validates stubbed response shapes. The suite uses those clients, not home-grown objects that merely happen to accept the same method names.
Ten tests cover submission fields, queued/running/completed transitions, failed jobs, a bounded wait, a throttled poll, private S3 reading, empty transcripts, access denial, invalid output and a zero-attempt configuration. The local Ruby 3.3.2 run records the precise SDK versions and results in the attachment. No audio left the machine and no recognition quality was measured.
Turn the example into a job lifecycle
For an application, store the input bucket/key, job name, language configuration, output bucket/key and current state together. Uploading the source object is a separate step. Confirm that the submitted object exists, that the region and permissions are correct, and that the eventual reader can retrieve the output.
A live smoke test should use a short consented clip with a known transcript. Verify submission, completion and authenticated retrieval independently, then exercise the behavior when a worker restarts between them. The offline suite cannot establish bucket policy, service permissions, KMS access, limits or costs.
Keeping the three operations separate gives each failure a place to land. A slow job can be polled later. A failed job can report its reason. An inaccessible result can remain an access error. None of those events needs to masquerade as a successful empty transcript or trigger a fresh transcription job.
Found a mistake or tried a different approach?
Send Alex a note ↗