NAME
Langertha::Role::Transcription - Role for APIs with transcription functionality
VERSION
version 0.503
transcription_model
The model name to use for transcription requests. Lazily defaults to default_transcription_model if the engine provides it, otherwise falls back to the general model attribute from Langertha::Role::Models.
transcription_file_part
my $part = $engine->transcription_file_part($audio, $filename);
Turns the audio argument of "transcription" into the multipart file part. $audio is one of:
a path to the audio file (a plain string);
a reference to a scalar holding the audio bytes:
\$bytes;an open filehandle, read to the end in binary mode;
a plain string that contains a NUL byte, taken as the audio bytes (no path can contain one). Prefer
\$bytes: a short content without a NUL would be taken as a path.
$filename is the name sent with the part; it defaults to the basename of the path, and to audio for in-memory content. Hosted APIs (OpenAI, Groq) detect the format from the filename's extension, so pass one such as speech.mp3 with in-memory audio.
transcription
my $request = $engine->transcription($audio, %extra);
my $request = $engine->transcription(\$bytes, filename => 'speech.mp3');
Builds and returns a transcription HTTP request object for the given audio: a path, \$bytes, or a filehandle (see "transcription_file_part"). filename in %extra sets the uploaded filename; the other %extra pairs are sent as form fields. Use "simple_transcription" to execute the request and get the transcript directly.
simple_transcription
my $text = $engine->simple_transcription($audio, %extra);
my $text = $engine->simple_transcription('/path/to/audio.mp3');
my $text = $engine->simple_transcription(\$audio_bytes,
filename => 'audio.mp3', language => 'en');
Sends a transcription request for the audio (a path, \$bytes or a filehandle, see "transcription_file_part") and returns the transcript text. Blocks until the request completes. Additional options such as language can be passed as %extra key/value pairs. "simple_transcription_f" is the non-blocking variant.
simple_transcription_f
my $text = await $engine->simple_transcription_f('/path/to/audio.mp3');
my $text = await $engine->simple_transcription_f(\$audio_bytes,
filename => 'audio.mp3', language => 'en');
Async variant of "simple_transcription": same arguments, returns a Future that resolves to the transcript text and fails with the same error text. The multipart upload goes through the engine's async backend (Langertha::Role::AsyncHTTP), so "user_agent_timeout" in Langertha::Role::HTTP bounds it on Net::Async::HTTP too; without that module it runs synchronously over LWP. The audio is read into the request body when the call is made.
simple_transcription_result
my $result = $engine->simple_transcription_result('/path/to/audio.mp3',
response_format => 'verbose_json',
'timestamp_granularities[]' => [qw( word segment )],
);
say $_->{word}, ' @ ', $_->{start} for @{ $result->{words} };
Like "simple_transcription", but returns the whole parsed answer as a HashRef (see "transcription_result" in Langertha::Role::OpenAICompatible) instead of only the text. verbose_json and word timestamps need a model that offers them (whisper-1, Groq, Whisper servers); OpenAI's default gpt-transcribe answers json only.
simple_transcription_result_f
my $result = await $engine->simple_transcription_result_f($audio,
response_format => 'verbose_json');
Async variant of "simple_transcription_result": returns a Future that resolves to the parsed answer as a HashRef, like "simple_transcription_f" does for the text.
simple_transcription_call
my $result = $engine->simple_transcription_call('/path/to/audio.mp3');
say $result->value; # the transcript text
say $result->usage->input_tokens if $result->has_usage;
my $segments = $result->raw->{segments}; # verbose_json
Like "simple_transcription", but returns a Langertha::CallResult: the transcript text as value, plus the provider's token usage (OpenAI's gpt-transcribe), this response's rate_limit, the model and the measured total_seconds. A JSON answer is kept whole in raw (segments, words, duration, a duration-billed usage).
It is not called simple_transcription_result because that name already returns the parsed answer as a HashRef, and keeps doing so.
simple_transcription_call_f
my $result = await $engine->simple_transcription_call_f(\$bytes,
filename => 'speech.mp3');
Async variant of "simple_transcription_call", sent like "simple_transcription_f": resolves to the Langertha::CallResult and fails with the same error text.
SEE ALSO
Langertha::Role::HTTP - HTTP transport layer
Langertha::Role::Models - Model selection (provides
transcription_model)Langertha::Engine::Whisper - Whisper-compatible transcription server
Langertha::Engine::Groq - Groq's hosted Whisper transcription
SUPPORT
Issues
Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.
IRC
Join #langertha on irc.perl.org or message Getty directly.
CONTRIBUTING
Contributions are welcome! Please fork the repository and submit a pull request.
AUTHOR
Torsten Raudssus <getty@cpan.org>
COPYRIGHT AND LICENSE
This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.
This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.