Product How it works Pricing Docs Log in
Rewriting the future of disability

The speech API for voices nobody else understands.

General speech recognition fails on dysarthric speech. SmfVox transcribes it — eleven per-language models, plus personal voice models you train on fifty read sentences.

Built by the Scott-Morgan Foundation. Serving people living with ALS, cerebral palsy, Parkinson’s and stroke.

$ curl https://dev.smfvox.ai/v1/transcribe \
    -H "Authorization: Bearer $KEY" \
    -F "file=@speech.wav" -F "model=smf_9f3ka27c"

{
  "text": "Het regende de hele middag zonder pauze.",
  "model": "smf_9f3ka27c",
  "confidence": 0.93,
  "billed_seconds": 6
}
A personal voice model, answering in about a second.
Product

Two endpoints nobody else sells.

Transcription against dysarthric baselines in eleven languages, and per-speaker adaptation as a service: post about fifty read sentences and get back a model tuned to that one voice.

Eleven dysarthric baselines

Czech, German, English, Spanish, French, Hungarian, Italian, Japanese, Dutch, Portuguese and Mandarin — per-language fine-tunes on dysarthric speech, not a general model with a language flag. Trained on a corpus of 67 clinical collaborations that took years of consent work to assemble.

Personal voice models

Fifty read sentences train an adapter for one speaker, in about a minute of GPU time. A new model goes live only if it beats the one in use on held-out sentences — quality cannot regress silently. Retrain as a voice changes; progressive conditions are the design case, not an edge case.

Consent built in, EU hosted

Deleting audio or a model is an API call, not a support ticket. Speech data lives on European infrastructure under pseudonymous speaker ids — leak one system and you have voices without names, leak the other and names without voices.

How it works

From key to personal model in an afternoon.

1

Get a key and transcribe

Sign up with Google, create an API key, and call the standard model for your language.

curl .../v1/transcribe -F "file=@a.wav" -F "language=nl"
2

Record the script

Download fifty sentences chosen for phonetic coverage, have the speaker read them, and upload each clip with its transcript.

GET /v1/training-transcripts?language=nl
3

Train and switch

One call starts training; the model id arrives by email minutes later. Point your transcription at it and every response uses that speaker’s own model.

POST /v1/train  →  "model": "smf_9f3ka27c"
Pricing

Metered for developers. Per seat for people.

Pay for what your integration uses. If you resell to end users, the licence requires per-seat pricing — a dysarthric speaker takes longer to say the same words, and nobody should pay more for being harder to understand.

Adapter training
€5 / training run

One flat fee per personal voice model trained.

  • About a minute of GPU per run
  • Failed runs are never charged
  • Promotion gate: quality cannot regress
  • Retrain any time as the voice changes
Read the walkthrough
Integrators
Let’s talk

Volume, per-seat resale licensing, and clinical deployments.

  • Per-enrolled-seat resale terms
  • Data-processing agreements
  • Priority language onboarding
Contact us
Documentation

Docs you can read before you sign up.

The full API reference and every guide are public. Copy-paste examples in curl, Python and JavaScript for all eleven endpoints.

API reference

Transcribe, models, training audio, runs, transcripts, usage and health — every request and response shape, with the recording scripts downloadable per language.

Open the reference →

Efficacy, stated exactly: on our first enrolled speaker, a personal adapter took character error rate from 22.6% to 15.8%, trained on twenty sentences with five held out.