Models are downloaded during setup, not bundled in the installer. Internet access, free disk space and Windows permission may be needed. This preview is not yet verified on a fresh PC.
01 / Your voice
Microphone → VRChat chat
Send translated speech to your chatbox without opening the keyboard.
Pauses new microphone capture when your VRChat mic is muted.
Can finish an already-captured sentence after you mute.
Optional cumulative translated messages while you speak.
02 / Their voices
VRChat audio → Subtitles
Read translated voices in a movable subtitle window. Start or close subtitles independently of your microphone translation.
Captures VRChat, rather than all Windows playback.
Choose the subtitle language separately.
Optional AI voice identification attaches speaker names to subtitles.
Toggle always on top.
Optional · Local AI
Recognize who’s speaking
A voice-identification model compares voices, not where people are standing, to help you follow speakers in a group.
On your PC
Identify voices
Matches clear speech to voice profiles and labels the subtitles. Voice identification can be turned off without stopping translation; turning it off unloads its model.
Your corrections
Save names and correct matches
Name a speaker or correct a mistaken match. Confirmed examples of an isolated voice help future matching. Saved names and voice profiles stay on your device; voice samples are not uploaded.
This identifies speakers; it does not separate voices talking over each other. Music, recordings and unclear speech can still cause mistaken matches.
First-time setup
Installation and setup
The app downloads and prepares its translation tools. This can take longer than the installation.
1
Install the app
Run the installer and choose Just me or All users. Open SayWhat? from Start when installation finishes.
2
Choose languages and prepare models
Follow Setup in the app. Review its download and Windows permission notice, then select Set up SayWhat?. If Windows asks for a restart, reopen the app and continue afterwards.
3
Open VRChat and start translation
Enable OSC in VRChat. In SayWhat?, start your voice, subtitles, or both. Completing setup does not start listening automatically.
↳
The app interface supports English, Simplified Chinese, Japanese and Korean. Interface language is separate from your translation languages.
Before you download
System requirements
Translation and VRChat share your GPU. Larger models use more memory and may increase delay.
Windows 11 x64
A Windows PC with VRChat and a microphone for your own voice. This is not a standalone Quest or iPad app.
WSL2 speech support
Setup prepares an app-owned Ubuntu speech engine. Windows virtualization, administrator permission and a restart may be required.
A compatible GPU
NVIDIA speech setup uses CUDA. Other GPUs need working Vulkan access in WSL; setup checks this and reports unsupported hardware instead of silently falling back to CPU.
Storage
The speech model is about 2.9 GB; the recommended translator is about 1.1 GB. Speech tools, build files and optional models need additional disk space.
Preview status: hardware compatibility, fresh-PC installation and long VRChat sessions are still being tested. A particular frame rate or translation delay is not guaranteed.
In-app downloads
Translation models
Download and switch models in the app. A larger model is not always faster or more accurate.
Default
Hy-MT2 1.8B
About 1.1 GB for Q4_K_M. Lower memory use leaves more GPU resources for VRChat.
Optional
Hy-MT2 7B
About 4.6 GB for Q4_K_M. An alternative for comparing translation quality, with higher memory use.
Optional
TranslateGemma 4B
About 2.5 GB for Q4_K_M. Requires a chosen spoken language and the built-in runtime. Gemma terms apply; the GGUF is a third-party conversion.
English ↔ Japanese only
Liquid AI LFM2 350M
About 229 MB for Q4_K_M. Other language pairs are blocked in the app. LFM Open License terms apply.
Use one shared translator, or choose separate models for your voice and subtitles in Advanced. Loading two models costs additional memory. Models have their own licenses and are downloaded from their listed third-party sources.
Privacy
Voice processing stays on your PC.
The local workflow does not upload your speech for recognition or translation. Saved people and voice fingerprints remain on your device; recognition is optional.
Setup downloads models and tools from third-party providers. Those providers receive normal download connection information, not microphone recordings.
Logs can contain spoken or translated text. Review them before exporting or posting them for support.
This website has no analytics scripts or tracking cookies. Hosting and download services may still receive normal request metadata.
FAQ
Help
Do I need LM Studio or FoxTrans separately?+
No separate translation app is required for the automatic setup. SayWhat? includes its capture host and translator runtime, then prepares speech support and downloads models after you approve setup. Existing compatible installations can be reused from Advanced.
Can my voice translate into several languages?+
Yes. Languages offers a primary language and up to two additional languages, off by default. Translations appear on separate lines in one chat bubble. They reuse one loaded translator, but take additional inference time and share the chatbox text limit. The selected model must support every language; Liquid's EN–JP model cannot translate to Chinese or Korean.
Will translated messages keep changing while I speak?+
Optional early messages are cumulative translations of the sentence so far. They are submitted chat bubbles, not keyboard drafts. The recognizer keeps the entire utterance as context and the completed sentence supplies the definitive translation. Updates are coalesced to roughly 1.5 seconds or slower; chatbox text is limited to 144 characters and 9 displayed lines. Long messages show a readable tail, not the whole paragraph at once.
Can it separate people who talk at the same time?+
Not as independent live audio streams. Optional voice recognition can help attach names to clear speech, but recognizing a person is different from separating overlapping voices. Overlap, music and recordings can still confuse recognition and translation.
Why is translation slow on a powerful PC?+
Delay includes the pause after speech, recognition finishing, translation and the chatbox send cadence. VRChat and local models also share GPU resources. Try the smaller translator, one shared model, or just one translation direction. Advanced has timing and troubleshooting information; a larger model alone does not solve every delay.
What happens when I close the app?+
SayWhat? asks whether to quit and release its model memory or keep running in the system tray. You can remember your choice and change it later in Settings. Closing the subtitle window stops subtitles independently.
Can I update or uninstall without losing models?+
The app checks for updates on launch unless you turn this off. Update now downloads and verifies the installer, then opens its update wizard; final confirmation and Windows permission remain visible. This path keeps your settings and models. The standalone installer also offers update, repair and uninstall with optional data choices. All-users operations keep each user’s private data.
How do I enable OSC in VRChat?+
Open VRChat’s Action Menu and select OSC → Enabled. SayWhat?'s sidebar checks local OSC availability and offers guidance if it cannot verify a connection. This is not a chat-delivery acknowledgment: chatbox visibility and user safety settings can still affect what other people see. See VRChat’s official OSC guide for current instructions.
Where do I report a problem?+
Open Advanced → Diagnostics in the app. Note which direction stopped, your selected models and the timing shown, then report the problem on GitHub. Review logs for private conversation text before sharing them.