NAME

Langertha::Role::Transcription - Role for APIs with transcription functionality

VERSION

version 0.503

transcription_model

The model name to use for transcription requests. Lazily defaults to default_transcription_model if the engine provides it, otherwise falls back to the general model attribute from Langertha::Role::Models.

transcription_file_part

my $part = $engine->transcription_file_part($audio, $filename);

Turns the audio argument of "transcription" into the multipart file part. $audio is one of:

  • a path to the audio file (a plain string);

  • a reference to a scalar holding the audio bytes: \$bytes;

  • an open filehandle, read to the end in binary mode;

  • a plain string that contains a NUL byte, taken as the audio bytes (no path can contain one). Prefer \$bytes: a short content without a NUL would be taken as a path.

$filename is the name sent with the part; it defaults to the basename of the path, and to audio for in-memory content. Hosted APIs (OpenAI, Groq) detect the format from the filename's extension, so pass one such as speech.mp3 with in-memory audio.

transcription

my $request = $engine->transcription($audio, %extra);
my $request = $engine->transcription(\$bytes, filename => 'speech.mp3');

Builds and returns a transcription HTTP request object for the given audio: a path, \$bytes, or a filehandle (see "transcription_file_part"). filename in %extra sets the uploaded filename; the other %extra pairs are sent as form fields. Use "simple_transcription" to execute the request and get the transcript directly.

simple_transcription

my $text = $engine->simple_transcription($audio, %extra);
my $text = $engine->simple_transcription('/path/to/audio.mp3');
my $text = $engine->simple_transcription(\$audio_bytes,
    filename => 'audio.mp3', language => 'en');

Sends a transcription request for the audio (a path, \$bytes or a filehandle, see "transcription_file_part") and returns the transcript text. Blocks until the request completes. Additional options such as language can be passed as %extra key/value pairs. "simple_transcription_f" is the non-blocking variant.

simple_transcription_f

my $text = await $engine->simple_transcription_f('/path/to/audio.mp3');
my $text = await $engine->simple_transcription_f(\$audio_bytes,
    filename => 'audio.mp3', language => 'en');

Async variant of "simple_transcription": same arguments, returns a Future that resolves to the transcript text and fails with the same error text. The multipart upload goes through the engine's async backend (Langertha::Role::AsyncHTTP), so "user_agent_timeout" in Langertha::Role::HTTP bounds it on Net::Async::HTTP too; without that module it runs synchronously over LWP. The audio is read into the request body when the call is made.

simple_transcription_result

my $result = $engine->simple_transcription_result('/path/to/audio.mp3',
    response_format => 'verbose_json',
    'timestamp_granularities[]' => [qw( word segment )],
);
say $_->{word}, ' @ ', $_->{start} for @{ $result->{words} };

Like "simple_transcription", but returns the whole parsed answer as a HashRef (see "transcription_result" in Langertha::Role::OpenAICompatible) instead of only the text. verbose_json and word timestamps need a model that offers them (whisper-1, Groq, Whisper servers); OpenAI's default gpt-transcribe answers json only.

simple_transcription_result_f

my $result = await $engine->simple_transcription_result_f($audio,
    response_format => 'verbose_json');

Async variant of "simple_transcription_result": returns a Future that resolves to the parsed answer as a HashRef, like "simple_transcription_f" does for the text.

simple_transcription_call

my $result = $engine->simple_transcription_call('/path/to/audio.mp3');
say $result->value;                                   # the transcript text
say $result->usage->input_tokens if $result->has_usage;
my $segments = $result->raw->{segments};              # verbose_json

Like "simple_transcription", but returns a Langertha::CallResult: the transcript text as value, plus the provider's token usage (OpenAI's gpt-transcribe), this response's rate_limit, the model and the measured total_seconds. A JSON answer is kept whole in raw (segments, words, duration, a duration-billed usage).

It is not called simple_transcription_result because that name already returns the parsed answer as a HashRef, and keeps doing so.

simple_transcription_call_f

my $result = await $engine->simple_transcription_call_f(\$bytes,
    filename => 'speech.mp3');

Async variant of "simple_transcription_call", sent like "simple_transcription_f": resolves to the Langertha::CallResult and fails with the same error text.

SEE ALSO

SUPPORT

Issues

Please report bugs and feature requests on GitHub at https://github.com/Getty/langertha/issues.

IRC

Join #langertha on irc.perl.org or message Getty directly.

CONTRIBUTING

Contributions are welcome! Please fork the repository and submit a pull request.

AUTHOR

Torsten Raudssus <getty@cpan.org>

COPYRIGHT AND LICENSE

This software is copyright (c) 2026 by Torsten Raudssus https://raudssus.de/.

This is free software; you can redistribute it and/or modify it under the same terms as the Perl 5 programming language system itself.