Smart speed and Voice boost
Smart speed (shorten the silences in narration) and Voice boost (compress and lift speech) are two device-wide settings, off by default, that each engine implements its own way. No wire change: both are client-side, and "time saved" is kept on the device.
| Android | iOS | Web | |
|---|---|---|---|
| Smart speed | every book: a silence-skipping AudioProcessor | not yet (why) | not offered |
| Voice boost | every book: a compressor/limiter AudioProcessor after Sonic | every book: an MTAudioProcessingTap per item | Web Audio compressor, trim and limiter, never in WebKit-only browsers, same-origin sources only |
The Voice boost preset
| Native (Android and iOS) | Web (below) | |
|---|---|---|
| High-pass | 80 Hz | - |
| Compressor | -20 dBFS, 3:1, 6 dB soft knee centred on the threshold, 10 ms attack, 200 ms release | the same numbers; the browser's knee runs from the threshold up |
| Make-up | +12 dB | the browser's own, plus a +1.6 dB trim: about +10 dB below the threshold |
| Limiter | -1 dBFS ceiling; attack about 1 ms (Android) or instant (iOS); release 80 ms (Android), 50 ms (iOS) | -3 dBFS, 20:1, 1 ms attack, 80 ms release (lands near -1 dBFS) |
The detector is per-sample peak, so a curve's unity point is threshold + makeup × ratio / (ratio - 1): -2 dBFS here, above narration's peaks, so the make-up is never taken back. Measured on real speech (active-speech RMS): quiet narration comes up about +10 to +11 dB, normal +8.5 dB, loud +5 dB, with peaks held at the ceiling.
The settings and the UI
src/stores/settings.ts:smartSpeedandvoiceBoost(PlaybackSettings), defaultfalse; a stored value counts only when it is exactlytrue. They reach the engine through the store'sconfigure(PlaybackConfig.smartSpeed/voiceBoost).src/playback/effects.tsholds the platform rules, pure:smartSpeedApplies(platform)(Android only) andsupportsVoiceBoost(env?)(Web Audio withcreateMediaElementSourceandcreateDynamicsCompressor, and notisWebKitOnly(userAgent): Safari and every browser on iOS/iPadOS, because routing an<audio>element through anAudioContextthere plays choppy audio, ignoresplaybackRateand suspends on the lock screen - WebKit bugs 240405, 311000, 261554).src/components/player/effects-model.tsturns those into what each switch says:smartSpeedRow(wheresmartSpeedAppliessays no: off and disabled,notInBrowseron the web,notOnIphoneon iOS; on Android the hint and the lifetime "Saved 2h 11m" oncehasSaved, a whole second) andvoiceBoostRow(disabled withnotInThisBrowserwhere the web can't run it).EffectsSettings(effects-settings.tsx) renders both switches as SettingsSettingRows. The speed sheet (SpeedSheet, under the presets) and Settings > Listening > Playback (PlaybackPane) mount the same component, so each setting lives in one place.- The full player's action row adds an effects pill (
sparkles) on tablet and desktop only (wide; the phone row is full):effectsPillsays "Saved 2h 11m" while Smart speed is on and has saved some, else "Voice boost" or "Smart speed" for whichever is on, nothing when both are off; the web and iOS only ever show Voice boost. It opens the speed sheet. - Strings live under the
effectsnamespace (effects.smartSpeed.*,effects.voiceBoost.*,effects.saved,effects.bookSaved,effects.pillLabel).
Android: one processor chain
effects/AudioEffects.kt is the seam the service and the module share. Its
renderersFactory(context) builds ExoPlayer's renderers with our own
AudiosiloAudioChain (effects/AudiosiloAudioChain.kt):
NarrationSilenceProcessor -> SonicAudioProcessor -> VoiceBoostProcessor
(book time) (speed) (what you hear)
Media3 1.5.1's DefaultAudioProcessorChain only accepts Media3's own (final) silence
processor, so the chain replaces it the way Pocket Casts does. Order matters: silence
skipping runs first, on the book's own frames, so its skipped-frame count is book time (what
getSkippedOutputFrameCount reports to the sink's position maths, and what time saved
counts); Sonic does the speed; Voice boost runs last, so its attack and release are real
time at any speed. One chain instance per sink.
NarrationSilenceProcessoris a Kotlin port of Media3 1.11.1'sSilenceSkippingAudioProcessor(Apache 2.0, provenance in the file header). It is vendored because 1.5.1's copy sizes its "maybe silence" buffer in frames but uses it as bytes (androidx/media#3271), keeping a quarter of the intended padding around each word in stereo, and fades per byte. Narration defaults: threshold 330 (about -40 dBFS peak), minimum silence 300 ms, retention ratio 0.25, at most 1 s of silence kept, mute to 10%. It also keeps a monotonic count of the book time the skipped frames were worth (savedUs; its per-flushskippedFrames, which the sink reads throughgetSkippedOutputFrameCount, still resets on every flush, asapplySkippingneeds).VoiceBoostProcessor: the native preset, then a final clip; one gain for every channel (stereo-linked). It is always active: changingisActivewould make the sink flush (a gap and a position jump), so the switch moves a ~20 ms ramp between dry and wet, and fully off the output is the input byte for byte.queueInputnever locks, and allocates nothing beyond the base class's reusable output buffer.- Switching: the module's
setConfigsends the custom session commandaudiosilo.SET_EFFECTS(extrassmartSpeed,voiceBoost); the service applies them (AudioEffects.apply:player.skipSilenceEnabled, which lets ExoPlayer drain and re-checkpoint around the change, and the boost's atomic flag) and saves them in theaudiosilo.playerSharedPreferences (effects.smart,effects.boost), so a service started without JS (the car, playback resumption) applies the listener's last choice. - Time saved:
AudioEffects.silenceSavedSeconds(process-wide, monotonic) rides on the module'sonProgressassilenceSaved. Silence-skip discontinuities emit no progress of their own.
Both processors have JVM tests (Testing).
iOS: a processing tap
Voice boost (VoiceBoostTap.swift, VoiceBoostDSP.swift)
An MTAudioProcessingTap on each AVPlayerItem's audio mix runs VoiceBoostDSP (pure Swift,
free of AVFoundation): the 80 Hz one-pole high-pass, a soft-knee compressor linked across
channels, the make-up gain and a peak limiter at the ceiling. It is our own arithmetic rather
than AUDynamicsProcessor/AUPeakLimiter, which need an AudioUnit graph pulling from the
tap, can't express the knee and ratio, and can allocate on a first render. Real-time rules
hold in process: no allocation, no locks, no Swift reference counting (the state is a plain
struct allocated in the tap's init and freed in finalize).
Every item gets its tap before it plays, switch on or off
(AudioEngine.attachVoiceBoostTaps): setting audioMix on a playing item rebuilds its
render chain, an audible ~1 s dropout on a device. So the mix is set while an item is still
waiting, for the current item and the two ahead (a window that slides on at each file
change, not the whole queue AVQueuePlayer holds: loadTracks
on a streamed asset reads the file's header over HTTP, and a book can be a hundred files; two
ahead means the next item has its tap before AVQueuePlayer prerolls it). Playback never waits
for a tap. The switch only flips one process-wide aligned 32-bit word that the tap reads once
per render call, and the DSP crossfades toward it over 20 ms; off and ramped down, the tap
leaves the audio untouched.
The tap sits on the item's audio mix, which AVPlayer runs in the item's own timeline, before the rate's time-pitch stage, so the attack and release are in book time.
VoiceBoostDSP has a host self-check that runs it on speech; it fails when speech comes up by
too little at any level (Testing).
Why Smart speed isn't on iPhone yet
An MTAudioProcessingTap can't drop frames, and changing AVPlayer's rate while it plays is
an audible dropout, so speeding through each silence stutters at every pause. A real version
needs a different audio path for downloaded books: an AVAudioEngine graph that can skip
frames. Until then smartSpeedApplies is Android only, the iOS engine accepts and ignores
config.smartSpeed, and its onProgress carries no silenceSaved.
Web: one Web Audio graph
service.web.ts implements Voice boost only:
- One
AudioContextand one chain to the destination: aDynamicsCompressorNode(VOICE_BOOST_COMPRESSOR), aGainNodetrim (VOICE_BOOST_TRIM_DB) and a second compressor as the limiter (VOICE_BOOST_LIMITER). ADynamicsCompressorNodehas no make-up setting and applies its own, so the trim and the limiter land the chain on the native curve above the knee;service.web.test.tsrecomputes the whole chain from the browser's curve. It is created lazily inside a gesture (the switch turned on, or a play tap with the setting on) and never torn down;ctx.resume()runs synchronously in that gesture. - Each
<audio>element gets oneMediaElementAudioSourceNodefor life (a secondcreateMediaElementSourcethrows), made only while the context runs (a suspended one would silence the book) and only for a same-origin source (sameOrigin,src/lib/same-origin.ts): a cross-origin element routed through Web Audio without CORS plays zeroes. So a book streamed from another signed-in server, or through the Metro dev server talking to a remote API, plays unboosted (a downloaded copy, served from the page's own origin under/_offline/, is boosted); a load of such a source onto an already-routed element swaps in a fresh element. - Switching off reconnects each source straight to the destination; nothing is rebuilt.
Time saved (src/playback/time-saved.ts)
"Time saved" is the book seconds Smart speed removed: the same as time saved at 1×, so speed isn't counted. Kept on the device, and counted outside React (only the two hooks at the end use it):
- The native bridge's
onSilenceSaved(total)(fromonProgress'ssilenceSaved) goes tonoteSilenceSaved(total, bookKey)with the playing book'scontentKey.silenceDeltacounts positive growth only: the first total this JS sees is a base (the Android service can outlive the JS, and its total then holds savings already counted), and a lower total only becomes the new base, adding nothing (a guard: Android's total is process-wide, so a rebuilt player never lowers it). persistedDocument('audiosilo.timeSaved')holds{ lifetime, books }. Never a write per tick: flushed when playback halts (the store'shaltAndPersist), when the app leaves the foreground, and every 30 s while it grows (anengineTicker, which Android's paused JS timers can't stop). Hydration merges what was counted before the read finished.onConnectionRemoveddrops that connection's books (withoutConnection); the lifetime total stays, since it is this device's.useTimeSaved()(the speed sheet's "Saved 2h 11m") anduseBookTimeSaved(cid, libraryId, path)(the book page's Smart speed saved row,listeningFigures'smartSpeedSaved).
Elsewhere: the engines and the native module are in Playback; CarPlay and Android Auto, which apply the same settings, are in Native integrations.