Building a Local Voice-to-Text and Rewriting Assistant on Linux

/recreational-programming

../../assets/common_files/thumbnails/voice-msg.webp

This article comes from a frustration I had with Discord Web.

I didn't have the possibility to directly send a voice message.

The only way to send one was to literally record my voice, save the file, and then send it manually.

But that was enough to give me an idea: if I had a way to simply speak and immediately send the damn message, that would already be great.

And then, imagine having a way to transcribe my voice directly into text and copy it straight to the clipboard; perfect for emails.

Even better: a way to reformulate what I've said in different tones!

That's the system I scripted yesterday, using keyboard shortcuts in my i3 environment, along with Bash, yad, notify, xclip, jq, ffmpeg, whisper.cpp, llama.cpp, systemd, and an open-source LLM model that I downloaded locally.

First, the features

I want to be able to press a keyboard shortcut that opens a small window with multiple buttons while immediately starting to record my voice.

The available actions are:

  • Stop recording and quit.

  • Stop recording and copy the audio file to the clipboard.

  • Stop recording, transcribe the audio to text, and copy the transcription to the clipboard.

  • Stop recording, transcribe the audio to text, reformulate it in a chosen tone, and copy the resulting text to the clipboard.

The setup

Downloads

On APT we download all the required packages we'll use with:

  
  

sudo apt update && sudo apt install -y \
    git \
    build-essential \
    cmake \
    ffmpeg \
    yad \
    xclip \
    curl \
    jq \
    libnotify-bin

  
  

whisper.cpp

First, we'll download it:

  
  

> git clone https://github.com/ggerganov/whisper.cpp.git

  
  

We enter the project:

  
  

> cd whisper.cpp

  
  

Now, we'll generate all the configurations build files directly in whisper/build:

  
  

whisper > cmake -S . -B build

  
  
  • -S . means the source directory is the one we are here (whisper)

  • -B build tells cmake to put the generated build files in build

But, what do we need this phase ?

One answer is because some code can take advantage of certain architecture-dependant feature such as:

  
  

AVX / AVX2
FMA
NEON
OpenMP
CUDA
Metal
BLAS

  
  

So it has to detect if it can use some compile flags.

Also, certain optional parts of the code can depend on libraries, so it check if they are installed, and where they are located.

Then, we finally compile it using the generated configuratiosn files:

  
  

whisper > cmake --build build -j

  
  

The -j flag means to build in parallel, using as much parallelism as the underlying build tool decides is available potentially using all CPU cores.

We could limit the dedicated cores for the build with:

  
  

whisper > cmake --build build -j 6

  
  

for example.

Now, the binaries are located in:

  
  

build/bin

  
  

We have:

  
  

bench                   libggml.so            libwhisper.so.1      test-parakeet-full-diffusion    test-whisper-zero-samples
libggml-base.so         libggml.so.0          libwhisper.so.1.9.5  test-parakeet-full-gb1          whisper-bench
libggml-base.so.0       libggml.so.0.26.0     main                 test-parakeet-full-jfk          whisper-cli
libggml-base.so.0.26.0  libparakeet.so        parakeet-cli         test-vad                        whisper-quantize
libggml-cpu.so          libparakeet.so.1      parakeet-quantize    test-vad-full                   whisper-server
libggml-cpu.so.0        libparakeet.so.1.9.5  test-common-utf8     test-whisper-buffer-loader      whisper-vad-speech-segments
libggml-cpu.so.0.26.0   libwhisper.so         test-parakeet        test-whisper-lang-detect-abort

  
  

First we just create a directory for the whisper model:

  
  

mkdir -p .local/share/whisper

  
  

Now, we want to download the model that will actually transcribe my voice to text:

  
  

whisper > ./models/download-ggml-model.sh .local/share/whisper/large-v3-turbo

  
  

Then, we create a systemd service (in systemd user space) to activate the whisper.cpp server to be able to receive the input file and to return the text:

  
  

mkdir -p .config/systemd/user

  
  

and we create this file whisper-server.service.

We write the following in it:

  
  

[Unit]
Description=Local whisper.cpp transcription server

[Service]
Type=simple
ExecStart=%h/whisper.cpp/build/bin/whisper-server \
    -m %h/.local/share/whisper/ggml-large-v3-turbo.bin \
    -t 12 \
    --host 127.0.0.1 \
    --port 8081

Restart=on-failure
RestartSec=3

[Install]
WantedBy=default.target

  
  

The %h part is a systemd specifier that expands to the home directory of the user running the service.

  
  

/home/juju

  
  

So:

  
  

ExecStart=%h/whisper.cpp/build/bin/whisper-server

  
  

is resolved by systemd as:

  
  

/home/juju/whisper.cpp/build/bin/whisper-server

  
  

Now, we reload the unit files known to the user instance of systemd, because we have just created a new service file and want systemd to take it into account:

  
  

systemctl --user daemon-reload

  
  

The --user flag tells systemctl to communicate with the systemd instance associated with the current user, rather than the system-wide instance.

This user instance searches for unit files in several user-specific locations, including:

  
  

~/.config/systemd/user/

  
  

It can also use system-wide directories containing user units, such as /etc/systemd/user/ and /usr/lib/systemd/user/.

llama.cpp

This is very similar to whisper.cpp setup.

We first download it:

  
  

> git clone https://github.com/ggerganov/llama.cpp.git

  
  

Enter into the project:

  
  

cd llama.cpp

  
  

Then configure the build:

  
  

llama.cpp > cmake -B build

  
  

Then compile:

  
  

llama.cpp > cmake --build build -j

  
  

Now, we'll download the LLM that will rewrite my transcribed text to something either more:

  • formal

  • warm

  • Concise

  • Rigorous

I chose Qwen3-8B with the Q4_K_M quantization. The 8B model is large enough to handle multilingual rewriting and instruction-following reliably, while the 4-bit quantization reduces the model to roughly 5 GB, making it practical to run locally on my hardware (my GTX 1660 Super has 6 GB of VRAM) .

The download link is:

https://huggingface.co

I downloaded it under ~/Téléchargements (~/Downloads).

Then, we'll create the service as .config/systemd/user/llama-server.service:

  
  

[Unit]
Description=Local llama.cpp server
After=network.target

[Service]
Type=simple
ExecStart=%h/llama.cpp/build/bin/llama-server -m %h/Téléchargements/Qwen3-8B-Q4_K_M.gguf -c 2048 -t 12 -ngl 99
Restart=on-failure
RestartSec=3

[Install]
WantedBy=default.target

  
  

Here, the server will listen on the port 8080 by default, so it does not conflict with the whisper server that listens to 8081.

And then we activate it:

  
  

systemctl --user daemon-reload
systemctl --user enable --now llama-server.service

  
  

Like that we could the following to request the model:

  
  

RESULT="$(
    curl -s http://127.0.0.1:8080/v1/chat/completions \
        -H 'Content-Type: application/json' \
        -d "$(
            jq -n \
                --arg text "$TEXT" \
                --arg style "$STYLE" \
                '{
                    messages: [
                        {
                            role: "system",
                            content: "You rewrite transcribed speech. Always preserve the original language of the input. Never translate the text into another language. Preserve the original meaning and information. Do not add new information. Return only the rewritten text."
                        },
                        {
                            role: "user",
                            content: ("Rewrite the following text in a " + $style + " style:\n\n" + $text + "\n\n/no_think")
                        }
                    ]
                }'
        )" |
    jq -r '.choices[0].message.content'
)"

  
  

As you can see, we build the JSON payload using jq.

Instead of manually constructing a Bash string such as "...", we let jq generate valid JSON for us.

Here, jq does not read any JSON document from stdin, which is why we use the -n (--null-input) flag. It tells jq that the input must be readen as an argument value.

The variables we want to inject into the JSON are passed as command-line arguments with --arg:

  
  

--arg text "$TEXT" 
--arg style "$STYLE" 

  
  

So jq directly creates the JSON payload internally.

Then, we know that llama-server returns a JSON response following the OpenAI-compatible chat completion format.

The generated responses are stored in an array named choices. Each element of this array represents one possible completion produced by the model.

A simplified response looks like this:

  
  

{
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": "WHISPER-OUTPUT"
      }
    }
  ]
}

  
  

The reason the model returns an array is that the API format is designed to support returning multiple completions for the same request. Even when only one response is generated, it is still wrapped in the choices array.

Here, we know that we just asked for one output, because we did not precise the variable n and its value, so it defaults to 1.

But, we could have passed it in the payload:

  
  

curl -sS http://127.0.0.1:8080/v1/chat/completions \
    -H 'Content-Type: application/json' \
    -d '{
        "n": 3,
        "messages": [
            {
                "role": "user",
                "content": "Rewrite this sentence in a warm tone."
            }
        ]
    }'

  
  

Btw, the equivalent HTML request looks like:

  
  

POST /v1/chat/completions HTTP/1.1
Host: 127.0.0.1:8080
Content-Type: application/json

{
  "messages": [
    {
      "role": "system",
      "content": "You rewrite transcribed speech..."
    },
    {
      "role": "user",
      "content": "Rewrite the following text..."
    }
  ]
}

  
  

We'll use the same curl method for the payload of whisper.cpp:

  
  

TEXT="$(
    curl -sS http://127.0.0.1:8081/inference \
        -H "Content-Type: multipart/form-data" \
        -F "file=@$FILE" \
        -F "response_format=json" \
        -F "language=auto" |
    jq -r '.text'
)"

  
  

Here, we put @$FILE to send the actual content of $FILE which is the .wav file containing the voice message, if we would just write $FILE, then only the literal filename would be sent.

Btw, here this is a multipart/form-data request, so the written HTML equivalent is:

  
  

<form
    action="http://127.0.0.1:8081/inference"
    method="post"
    enctype="multipart/form-data"
>
    <input type="file" name="file">

    <input
        type="hidden"
        name="response_format"
        value="json"
    >

    <input
        type="hidden"
        name="language"
        value="auto"
    >

    <button type="submit">Transcribe</button>
</form>

  
  

The bash script

We'll create the following file:

  
  

.local/bin/voice-recorder

  
  

And write the following code in it:

  
  

#!/usr/bin/env bash

DIR="$HOME/Recordings"
mkdir -p "$DIR"

FILE="$DIR/voice-$(date +%Y%m%d-%H%M%S).wav"

ffmpeg \
    -f pulse \
    -i default \
    -ar 16000 \
    -ac 1 \
    -c:a pcm_s16le \
    "$FILE" &

PID=$!

yad \
    --title="Voice recorder" \
    --text="🎙 Recording…" \
    --button="Copy audio:0" \
    --button="Transcribe:1" \
    --button="Stop:3" \
    --button="Formal:4" \
    --button="Warm:5" \
    --button="Concise:6" \
    --button="Rigorous:7" \
    --width=320 \
    --height=120

CHOICE=$?

kill -INT "$PID"
wait "$PID" 2>/dev/null

if [ "$CHOICE" -eq 0 ]; then
    printf 'file://%s\n' "$FILE" |
        xclip -selection clipboard -t text/uri-list

    notify-send \
        "Voice recorder" \
        "Audio copied to clipboard"

elif [ "$CHOICE" -eq 1 ]; then

    TEXT="$(
        curl -sS http://127.0.0.1:8081/inference \
            -H "Content-Type: multipart/form-data" \
            -F "file=@$FILE" \
            -F "response_format=json" \
            -F "language=auto" |
        jq -r '.text'
    )"

    printf '%s' "$TEXT" |
        xclip -selection clipboard

    notify-send \
        "Voice transcription" \
        "Transcription copied to clipboard"
elif [ "$CHOICE" -eq 3 ]; then
    exit 0
elif [ "$CHOICE" -ge 4 ] && [ "$CHOICE" -le 7 ]; then
   
    TEXT="$(
        curl -sS http://127.0.0.1:8081/inference \
            -H "Content-Type: multipart/form-data" \
            -F "file=@$FILE" \
            -F "response_format=json" \
            -F "language=auto" |
        jq -r '.text'
    )"

    case "$CHOICE" in
        4)
            STYLE="formal and professional"
            ;;
        5)
            STYLE="warm, friendly and natural"
            ;;
        6)
            STYLE="concise and direct"
            ;;
        7)
            STYLE="rigorous, precise and well-structured"
            ;;
    esac

    RESULT="$(
        curl -s http://127.0.0.1:8080/v1/chat/completions \
            -H 'Content-Type: application/json' \
            -d "$(
                jq -n \
                    --arg text "$TEXT" \
                    --arg style "$STYLE" \
                    '{
                        messages: [
                            {
                                role: "system",
                                content: "You rewrite transcribed speech. Always preserve the original language of the input. Never translate the text into another language. Preserve the original meaning and information. Do not add new information. Return only the rewritten text."
                            },
                            {
                                role: "user",
                                content: ("Rewrite the following text in a " + $style + " style:\n\n" + $text + "\n\n/no_think")
                            }
                        ]
                    }'
            )" |
        jq -r '.choices[0].message.content'
    )"

    printf '%s' "$RESULT" |
        xclip -selection clipboard

    notify-send \
        "Voice transcription" \
        "$STYLE version copied to clipboard"

fi


  
  

This is a very simple script, the only part that is worth mentioning is maybe this part:

  
  

ffmpeg \
    -f pulse \
    -i default \
    -ar 16000 \
    -ac 1 \
    -c:a pcm_s16le \
    "$FILE" &

PID=$!

yad \
    --title="Voice recorder" \
    --text="🎙 Recording…" \
    --button="Copy audio:0" \
    --button="Transcribe:1" \
    --button="Stop:3" \
    --button="Formal:4" \
    --button="Warm:5" \
    --button="Concise:6" \
    --button="Rigorous:7" \
    --width=320 \
    --height=120

CHOICE=$?

kill -INT "$PID"
wait "$PID" 2>/dev/null

  
  

We launch ffmpeg to record the voice in the background with the &, then we get its PID, launch the yad window to block the script and give the user the choice when to stop the record and what to do with it.

Then, we get the choices (yad button number), send a termination signal (SIGINT) to the ffmpeg process.

But:

  
  

kill -INT "$PID"

  
  

Does not wait the process to be terminated, it just sends the SIGINT, so we have to wait for it to be terminated, that's why we do the following directly afterward:

  
  

wait "$PID"

  
  

In fact we also hide stderr messages with the 2>/dev/null.

We have 3 descriptors:

  
  

0 = stdin 
1 = stdout  
2 = stderr  

  
  

In everyday bash scripts, we rarely see the 0 descriptor being used.

Indeed, something like:

  
  

some-command 0>/dev/null

  
  

or equivalently:

  
  

some-command </dev/null

  
  

It is used when we want to make sure a command cannot read keyboard input. Its stdin is redirected from /dev/null, so any attempt to read from standard input immediately encounters EOF instead of waiting for user input which terminates its reading.

Now, we just make it executable with:

  
  

chmod +x .local/bin/voice-recorder

  
  

The i3 setup

Now, we just have to add a keyboard shortcut entry to the i3 configuration, I've chosen Ctrl+Shift+v.

So inside .config/i3/config, I add:

  
  

bindsym $mod+Shift+v exec --no-startup-id ~/.local/bin/voice-recorder

  
  

--no-startup-id tells i3 to run the command, but not to use i3’s startup-notification tracking for it.

The “startup notification” is an X11 desktop protocol/mechanism used by launchers and window managers to track the fact that an application is in the process of starting.

Roughly:

  
  

i3 launches application
 |
 V
a startup ID is created
 |
 V
that ID is passed through the environment
 |
 V
the application/window announces itself with that startup ID
 |
 V
i3 knows:

“this window belongs to that launch request”

  
  

This helps desktop environments with things like associating a new window with the launcher action, showing a busy cursor, and knowing when application startup has completed.

We can check the configuration synthax:

  
  

i3 -C ~/.config/i3/config

  
  

Finally, we just have to reload i3 config without restarting i3:

  
  

i3-msg reload

  
  

Conclusion

Hope this article was usefull ;)