Windows app · Preview

Voice translation
for VRChat

Translate your voice into VRChat chat and other players’ voices into subtitles. Optional local AI identifies speakers in group conversations.

Windows 11 x64 · Runs locally

Example conversationIllustration
Your voiceEnglish → Chinese

“Are we meeting here tomorrow?”

明天我们在这里见面吗?

VRChat chat
Their voicesChinese → English

“这边比较安静。”

Ari Example speaker

It’s quieter over here.

Example messages. Choose a separate language for each direction.

Windows download

Install SayWhat?

The public installer is being prepared. Downloads are not available yet.

Preview · Windows 11 x64 · Unsigned installer

Installer being prepared

Models are downloaded during setup, not bundled in the installer. Internet access, free disk space and Windows permission may be needed. This preview is not yet verified on a fresh PC.

01 / Your voice

Microphone → VRChat chat

Send translated speech to your chatbox without opening the keyboard.

  • Pauses new microphone capture when your VRChat mic is muted.
  • Can finish an already-captured sentence after you mute.
  • Optional cumulative translated messages while you speak.
02 / Their voices

VRChat audio → Subtitles

Read translated voices in a movable subtitle window. Start or close subtitles independently of your microphone translation.

  • Captures VRChat, rather than all Windows playback.
  • Choose the subtitle language separately.
  • Optional AI voice identification attaches speaker names to subtitles.
  • Toggle always on top.

Optional · Local AI

Recognize who’s speaking

A voice-identification model compares voices, not where people are standing, to help you follow speakers in a group.

On your PC

Identify voices

Matches clear speech to voice profiles and labels the subtitles. Voice identification can be turned off without stopping translation; turning it off unloads its model.

Your corrections

Save names and correct matches

Name a speaker or correct a mistaken match. Confirmed examples of an isolated voice help future matching. Saved names and voice profiles stay on your device; voice samples are not uploaded.

This identifies speakers; it does not separate voices talking over each other. Music, recordings and unclear speech can still cause mistaken matches.

First-time setup

Installation and setup

The app downloads and prepares its translation tools. This can take longer than the installation.

  1. 1

    Install the app

    Run the installer and choose Just me or All users. Open SayWhat? from Start when installation finishes.

  2. 2

    Choose languages and prepare models

    Follow Setup in the app. Review its download and Windows permission notice, then select Set up SayWhat?. If Windows asks for a restart, reopen the app and continue afterwards.

  3. 3

    Open VRChat and start translation

    Enable OSC in VRChat. In SayWhat?, start your voice, subtitles, or both. Completing setup does not start listening automatically.

The app interface supports English, Simplified Chinese, Japanese and Korean. Interface language is separate from your translation languages.

Before you download

System requirements

Translation and VRChat share your GPU. Larger models use more memory and may increase delay.

Windows 11 x64

A Windows PC with VRChat and a microphone for your own voice. This is not a standalone Quest or iPad app.

WSL2 speech support

Setup prepares an app-owned Ubuntu speech engine. Windows virtualization, administrator permission and a restart may be required.

A compatible GPU

NVIDIA speech setup uses CUDA. Other GPUs need working Vulkan access in WSL; setup checks this and reports unsupported hardware instead of silently falling back to CPU.

Storage

The speech model is about 2.9 GB; the recommended translator is about 1.1 GB. Speech tools, build files and optional models need additional disk space.

Preview status: hardware compatibility, fresh-PC installation and long VRChat sessions are still being tested. A particular frame rate or translation delay is not guaranteed.

In-app downloads

Translation models

Download and switch models in the app. A larger model is not always faster or more accurate.

Default

Hy-MT2 1.8B

About 1.1 GB for Q4_K_M. Lower memory use leaves more GPU resources for VRChat.

Optional

Hy-MT2 7B

About 4.6 GB for Q4_K_M. An alternative for comparing translation quality, with higher memory use.

Optional

TranslateGemma 4B

About 2.5 GB for Q4_K_M. Requires a chosen spoken language and the built-in runtime. Gemma terms apply; the GGUF is a third-party conversion.

English ↔ Japanese only

Liquid AI LFM2 350M

About 229 MB for Q4_K_M. Other language pairs are blocked in the app. LFM Open License terms apply.

Use one shared translator, or choose separate models for your voice and subtitles in Advanced. Loading two models costs additional memory. Models have their own licenses and are downloaded from their listed third-party sources.

Privacy

Voice processing stays on your PC.

The local workflow does not upload your speech for recognition or translation. Saved people and voice fingerprints remain on your device; recognition is optional.

Setup downloads models and tools from third-party providers. Those providers receive normal download connection information, not microphone recordings.

Logs can contain spoken or translated text. Review them before exporting or posting them for support.

This website has no analytics scripts or tracking cookies. Hosting and download services may still receive normal request metadata.

FAQ

Help

Do I need LM Studio or FoxTrans separately?

No separate translation app is required for the automatic setup. SayWhat? includes its capture host and translator runtime, then prepares speech support and downloads models after you approve setup. Existing compatible installations can be reused from Advanced.

Can my voice translate into several languages?

Yes. Languages offers a primary language and up to two additional languages, off by default. Translations appear on separate lines in one chat bubble. They reuse one loaded translator, but take additional inference time and share the chatbox text limit. The selected model must support every language; Liquid's EN–JP model cannot translate to Chinese or Korean.

Will translated messages keep changing while I speak?

Optional early messages are cumulative translations of the sentence so far. They are submitted chat bubbles, not keyboard drafts. The recognizer keeps the entire utterance as context and the completed sentence supplies the definitive translation. Updates are coalesced to roughly 1.5 seconds or slower; chatbox text is limited to 144 characters and 9 displayed lines. Long messages show a readable tail, not the whole paragraph at once.

Can it separate people who talk at the same time?

Not as independent live audio streams. Optional voice recognition can help attach names to clear speech, but recognizing a person is different from separating overlapping voices. Overlap, music and recordings can still confuse recognition and translation.

Why is translation slow on a powerful PC?

Delay includes the pause after speech, recognition finishing, translation and the chatbox send cadence. VRChat and local models also share GPU resources. Try the smaller translator, one shared model, or just one translation direction. Advanced has timing and troubleshooting information; a larger model alone does not solve every delay.

What happens when I close the app?

SayWhat? asks whether to quit and release its model memory or keep running in the system tray. You can remember your choice and change it later in Settings. Closing the subtitle window stops subtitles independently.

Can I update or uninstall without losing models?

The app checks for updates on launch unless you turn this off. Update now downloads and verifies the installer, then opens its update wizard; final confirmation and Windows permission remain visible. This path keeps your settings and models. The standalone installer also offers update, repair and uninstall with optional data choices. All-users operations keep each user’s private data.

How do I enable OSC in VRChat?

Open VRChat’s Action Menu and select OSC → Enabled. SayWhat?'s sidebar checks local OSC availability and offers guidance if it cannot verify a connection. This is not a chat-delivery acknowledgment: chatbox visibility and user safety settings can still affect what other people see. See VRChat’s official OSC guide for current instructions.

Where do I report a problem?

Open Advanced → Diagnostics in the app. Note which direction stopped, your selected models and the timing shown, then report the problem on GitHub. Review logs for private conversation text before sharing them.