For speech-language pathologists and AAC specialists
Clinical information, including the parts that count against us
This page exists so that you can screen this tool in ten minutes and reach a defensible decision. It is written on the assumption that you will look for the authorship problem first, because that is the correct thing to look for.
What the tool does, precisely
One flat screen. Four large word tiles, drawn as a prefix of a single ranked 42-word adult starter list, with words 7 to 42 reached by typing one letter into the system keyboard, which filters the grid live. No categories, no folders, no paging, no nested location levels.
The user selects two to four words and taps one primary button. Apple's on-device Foundation Models return up to three candidate sentences, which are shown progressively as each clears the deterministic guards. Beneath them, in every state, sits a literal row that speaks the selected words verbatim and never touches the model. The user taps one. Only then is anything spoken.
Word tiles, saved phrases, typing, all voices, speech during calls, backup and export work identically whether or not the device supports Apple Intelligence. On ineligible hardware a deterministic rule-based expander produces up to two options plus the literal row, so the interaction has the same shape, the same tap count and the same confirmation step on every supported device.
Authorship, and how this differs from Facilitated Communication
The objection is legitimate and we will state it for you. A model that turns water cold now into a fluent request is producing words the person did not select. If the model is the author, the utterance is not the person's communication. ASHA's position on Facilitated Communication is unambiguous: it is a discredited technique, and information obtained through it should not be considered the communication of the person with a disability, because messages are authored by the facilitator.
Six structural properties, none of which is a setting a user or a partner can turn off:
-
No physical support, no partner in the selection loop
Selection is by direct touch, or by the person's own iOS access method. There is no facilitator hand, arm or shoulder, and the primary flow never allows a partner to operate the expansion on the person's behalf.
-
Nothing is ever spoken without an explicit tap
Automatic speech of generated text does not exist in the product. There is no auto-speak setting, no timeout that speaks, and no configuration that changes this.
-
Multiple candidates, not one
Offering two or three readings makes the person a selector, which is the field's own model of authorship, rather than a passive recipient of a single interpretation. This was not a design flourish. In development testing, a single-candidate design invented a wish from coffee, too cold and invented a fact from card, blocked, not me. Candidate selection removes that failure mode structurally rather than by prompt engineering.
-
A permanent literal bypass
Say my words is present on every candidate screen in every state, including model failure, and speaks the selected words verbatim. The whole language model can also be switched off in Settings, in which case the deterministic expander is used instead.
-
An exportable authorship trail
When logging is enabled, every record stores the raw selected words alongside the final spoken utterance, plus flags for whether a suggestion was accepted and whether it was edited. You can demonstrate authorship from the record rather than assert it. Logging is off by default and is switched on deliberately, by the person or by their caregiver.
-
Guards that refuse novel content
Two deterministic filters run between the model and the screen. A polarity guard rejects any candidate whose negation, refusal or affirmation does not match the selected words — so tired, not sad can never surface as I am sad. A content guard locks a guarded lexicon, including body parts, medicines and people, so a candidate may not introduce a guarded term the person did not select, and caps novel content generally. A candidate that fails either guard is discarded before it is displayed, not after.
The honest residual: the surface wording of an accepted candidate is the model's, not the person's. That is exactly why the candidate is shown at large size and must be tapped, why there are alternatives, why the literal row is permanent, and why the raw selection is preserved in the log. If those safeguards are not sufficient for a given client, the deterministic expander and the literal row remain available with the model switched off entirely.
Feature match
Organised by the clinical feature categories from Gosnell, Costello and Shane (2011), which is the framework most app feature matches are built on.
| Category | What this tool does |
|---|---|
| Purpose | Expressive communication aid for literate adults with acquired speech loss. Text-first. Not a therapy tool and not a language-learning system. |
| Language representation | Single words, spelling through the system keyboard, pre-stored phrases, and sentence expansion from selected words. No symbol set is used at all in version 1. |
| Vocabulary | 42-word ranked adult starter list plus a 600-word filtered lexicon, and 108 pre-written phrases in nine packs, covering hospital, everyday, feelings, self-advocacy, repair, negation, people and emergency. Emotional and safeguarding vocabulary is in the core set, not an add-on. Fully editable, and the person's own phrases always outrank generated ones. |
| Speech output | Apple system voices with novelty voices filtered out, Eloquence voices as a first-class category, and Apple Personal Voice. Five labelled rate steps. Voice is stored by identifier and re-verified at every launch. Speech during phone and FaceTime calls. |
| Display | Portrait, one flat screen, four tiles, no scrolling grid. Primary action bottom-anchored and at least 120 pt tall. Dynamic Type from Large to AX5 on every screen; columns reduce, type never shrinks. Spoken sentence displayed at 34 pt or larger, with word-level highlight synchronised to speech. Full-screen Show my words mode auto-fits 96 pt down to 48 pt, keeps the screen awake, and flips 180 degrees for a partner across a table. |
| Feedback | Word-level speech highlight, three-cue state changes (border weight, ink inversion, label change) so no state is signalled by colour alone, and a reachable Stop on every screen that can produce audio. |
| Rate enhancement | Sentence expansion from two to four selected words, pre-stored phrases, system word prediction and autocorrect in the keyboard, and a one-tap save of any spoken sentence into the phrase bank. |
| Access | Direct touch, VoiceOver with Direct Touch on the grid, Switch Control, Full Keyboard Access with ten keyboard shortcuts, Voice Control with per-tile synonym labels, and iOS eye and head tracking. No gesture is ever the only route to an action. |
| Motor and sensory support | Minimum 88 pt targets with a 64 pt absolute floor, minimum 24 pt gaps, no timed interactions of any kind, no double-tap requirement, and single-thumb reachability verified against a one-handed grip model on a 375 pt-wide device. |
| Support and stability | Named support contact, a public changelog, local backup and export to a single file through the share sheet, and a stated data-export guarantee. Voice, layout and app icon are never changed in an update. |
| Miscellaneous | Editing lock with a four-digit code, Guided Access restrictions, Assistive Access support, no in-app purchases of any kind, and no account, login or network permission. |
Access methods
Every route below is a first-class path, not an accident of the platform. Each is verified as a release gate rather than assumed.
- Direct touch — minimum 88 pt targets, 64 pt absolute floor, minimum 24 pt separation.
- VoiceOver — every control is a real button with a label, a hint and custom actions. Direct Touch is enabled on the word grid so the person is not forced to swipe through tiles.
- Switch Control — every action reachable by scanning, with grouped rows in explicit frequency order so common utterances are reached in fewer scan steps.
- Full Keyboard Access — exactly ten keyboard shortcuts, including stop.
- Voice Control — every tile carries several spoken synonyms, and no two tiles share one.
- Eye and head tracking — supported through the iOS system features rather than a bespoke implementation, with target sizes chosen accordingly.
- Dynamic Type Large to AX5 on every screen, plus Reduce Motion, Show Button Shapes, Increase Contrast and Bold Text.
Where a person needs mounted, hands-free or alternative access as their primary method throughout the day, a dedicated device is the correct recommendation and this tool is not a substitute for it. See what this tool is not for.
The trial report
Logging is off by default. When you and the person switch it on, the app captures an event stream rather than aggregates, in the tradition of Language Activity Monitoring, and generates a report on request. The report is free, like everything else in the app: a clinician should never have to buy the evidence.
Automatically captured
- Trial days, sessions, total utterances, utterances per day
- Mean utterance length in words
- Average communication rate and peak communication rate, in words per minute
- Keystroke savings, as a percentage
- Language representation method mix: spelling, word selection, phrase bank, sentence expansion
- Spontaneous compared with pre-stored ratio
- Core compared with extended vocabulary frequency
- Corrections per utterance
- Unique word count and type-token ratio
- Suggestion accept rate and suggestion edit rate — the authorship evidence
- Font size and target size actually in use, for the vision section
- Timestamped sample utterances, verbatim, each paired with the raw words the person selected
Clinician-authored, and clearly marked as such
The report structure follows the elements a funding claim is normally built from: current communication impairment and severity; whether natural modes can meet daily needs, with a rule-out rationale for speech, writing, gesture and low-tech boards; functional communication goals and treatment options; the rationale for tool and feature selection; demonstrated ability to use the tool; cognitive and physical abilities; the treatment and training plan; and a disclosure and signature block.
Exports
PDF report, CSV event stream, and a SALT-compatible plain-text transcript. Everything leaves the device through the iOS share sheet, under the person's control. Nothing is uploaded anywhere by the app, because the app cannot reach a network.
The disclaimer that ships in the report footer
You may keep or delete it. It reads: this software is speech generating software, compare HCPCS E2511. It is not a dedicated speech generating device and does not meet Medicare durable medical equipment criteria when installed on a general computing device, per local coverage article A52469. This report documents a functional trial and is not a substitute for a comprehensive AAC evaluation.
How to run a two-week trial
-
Install the app — there is nothing to request
The app is free for everyone, clinicians included, with everything unlocked. Install it from the App Store on your own device and on the client's. No codes, no license request, no account.
-
Set up in one session, then stop
Enter the person's name, the two or three people they most need to name, and their doctor, in About me. Add one or two phrase packs. Do not build a custom board. The default content is designed to work with no configuration at all, and setup burden is one of the best-documented causes of abandonment.
-
Switch on logging, with consent, and say what it records
It records what was tapped, what was spoken and when. It records nothing else, and it goes nowhere. Show the person the export before you begin, so consent is informed.
-
Set the editing lock
A four-digit code, plus Guided Access if the device is shared on a ward. This prevents accidental deletion of the vocabulary, which is a common and severe failure.
-
Train the partner as well as the person
Wait time, augmented input, verification, and how to hand the phone over. Partner training is the most-cited factor in whether a system is still in use at three months.
-
Use it in at least three environments
Clinic, home, and one environment with an unfamiliar partner — a pharmacy counter, a reception desk, a phone call. Tag them. Environment coverage is what a funding reviewer reads for.
-
Code prompt level per utterance
The log accepts a per-utterance prompt or cue code and a three-point adequacy rating. Independent use is what funding decisions turn on, so code it as you go rather than reconstructing it later.
-
Export at the end and edit freely
The PDF is a starting document, not a submission. Everything auto-filled is marked as auto-filled. Delete anything you do not stand behind.
Evidence, and its limits
Plainly: there is no efficacy study for this tool. It has not been peer reviewed, it is not clinically validated, and it has not been shown to improve any outcome. Any claim otherwise would be false, and you should treat any AAC app that makes one with suspicion.
What does exist is measured engineering evidence, from development testing of the on-device model on Apple silicon on 2 August 2026:
| Condition | Average | Worst observed |
|---|---|---|
| First generation from cold, with no warm-up | — | 4.20 s |
| Single sentence, once warm | 0.34 s | 0.36 s |
| Three candidates, once warm | 0.66 s | 0.80 s |
Two consequences were designed in as a result. A discarded warm-up generation runs at launch, because a 4.2-second first utterance is unacceptable at a bedside. And three candidates cost roughly 0.3 s more than one, which is why the safer interaction is also the shipped one.
The failure findings matter more than the latency. At a moderate temperature with a plausible single-candidate prompt, three of six test outputs were unusable: an invented wish, an invented fact, and one word-salad. Tightening the prompt and lowering the temperature removed the invention but collapsed the output into telegraphic echo — pain, left leg, worse at night became Pain in left leg worse at night, which is barely more than reading the tapped words aloud. Neither single-candidate design was shippable. The three-candidate design is the resolution of that specific failure, and it is documented here rather than buried because it is the strongest argument for the safeguards in the section above.
Two known residual defects, with their mitigations: the model occasionally returns duplicate candidates, so candidates are deduplicated and backfilled to three; and possessives can be resolved to the wrong person, so personal vocabulary entered in About me is injected into the prompt, which turns please call your daughter into please call Sarah. Personal vocabulary is therefore a requirement, not a nicety.
An automated regression harness runs a fixed corpus against the model with pass and fail gates before every release, and the app carries an on-device switch that disables generation if a model update degrades guard behaviour in the field.
Regulatory and funding position
It is not a medical device
It is a general communication aid in the same family as a letter board. It makes no claim to diagnose, treat, cure or mitigate any disease, and no claim to restore or improve any bodily function. Nothing on this website or in the app store listing says otherwise, and that discipline is deliberate: marketing is the entire risk surface for a product like this.
It is not a dedicated speech generating device
Medicare local coverage article A52469 states that desktop, laptop, tablet, smartphone and other general computing devices are not durable medical equipment, and that Medicare will reimburse for speech generating software only, HCPCS code E2511, when installed on a general computing device. The app is exactly that: software on a general computing device. It is not a dedicated device and does not claim to be.
It does not undercut a funded claim
A cash-purchased app is not a billed device and does not create a duplicate claim. Where a person needs alternative access, mounting, ruggedness or all-day battery, those needs are documented on their own terms. Many clinicians use a phone-based aid explicitly as a bridge while a funding claim is pending, and as a second always-with-you method afterwards. We will not give you language that overstates what an app can do, because the first reviewer who notices costs you the claim.
Privacy for institutional use
The app has no network entitlement, uses no third-party code, and collects nothing. Sentences are generated by Apple's on-device models and speech uses Apple's on-device synthesiser. No utterance is transmitted anywhere, so there is no processor to place under an agreement. The privacy policy is written to be handed to a compliance officer without translation.
Free for everyone, and the no-commission policy
Free for everyone
The app is free on the App Store, for clinicians and clients alike, with everything included. There are no license codes, because there is nothing to unlock, and there is no lite build. A feature match cannot be performed on a crippled version, and being asked to do one is a legitimate reason to reject an app.
Questions about clinical use: singhanhad78@gmail.com.
Zero commissions, permanently
No commissions, no referral fees, no affiliate revenue, no sponsored posts, no paid placements, ever. This is published rather than merely intended, because Medicare coverage determination L33739 requires that the evaluating clinician has no financial relationship with the supplier, and because ASHA requires disclosure of financial relationships with device manufacturers.
Because the app is free for everyone, a clinician who reviews it publicly has received nothing that the public does not receive.
What this tool is not for
- People who cannot read. Version 1 has no symbol set at all. For a client with significant alexia, a scene-based or photographic system chosen with you is the right tool, and this is not it.
- Primary access by eye gaze or switch scanning all day, mounted. iOS access methods are supported, but a phone is not a mounted, ruggedised, all-day device, and pretending otherwise would damage your funding case rather than help it.
- Therapy. There are no drills, no naming tasks and no exercises, deliberately. Adults with acquired aphasia consistently report resenting being quizzed by a tool they wanted to talk with.
- Children. The vocabulary, the register and the visual design are adult throughout.
- Android, Windows or web. iPhone and iPad only, because the sentence engine and the voice stack are Apple on-device frameworks.
- Anyone who wants a subscription-funded roadmap. There is no subscription and no purchase — the app is free — so there is no revenue behind it at all. That is a deliberate trade and you should know it when you assess how long the tool is likely to be maintained. Local backup and export exist so that nobody's vocabulary is trapped if that trade goes badly.
Sources
- CMS local coverage article A52469 — speech generating devices policy article, including the general computing device exclusion and the E2511 coding guideline.
- CMS local coverage determination L33739 — coverage criteria, including the evaluating clinician's financial relationship criterion.
- ASHA: Medicare speech generating device policy.
- ASHA Practice Portal: Aphasia — assessment components, communicative competencies, and the abandonment figure of approximately one third of cases.
- Gosnell, Costello and Shane (2011), Using a clinical approach to answer “What communication apps should we use?” — the eleven feature categories used in the table above.
- Hill and Romich, Language Activity Monitoring and the AAC Performance Report Tool — the event-stream logging and summary measures the report follows.
- Apple: Foundation Models framework — the on-device generation used, including guided generation.
- Apple Support: Personal Voice — including the rule that third-party apps may speak with a Personal Voice but may not capture its audio.
- Latency, invention and candidate results: development testing on Apple silicon, 2 August 2026. Method and raw outputs are available on request.