FAQ
Is this Auto-Tune? No. Auto-Tune is Antares' product and trademark; it defined this entire category, it's legendary, and it's worth checking out. This is a voice tuner: it does a similar job, but it is entirely custom DSP from the ground up, with its own character and its own decisions, and these pages explain every one of them.
Why does correction start a moment after a note does? Because the tuner waits (about 15 ms) for the pitch estimate to stop moving before it acts. Detectors guess badly at note onsets, often a full octave flat, and correcting a bad guess sounds like a scoop up into every phrase. During the wait your voice passes through dry; 15 ms is shorter than any note you can sing, and bends inside a note never re-trigger it. The full story is on the deciding page.
Why don't I get the chipmunk effect? Pitch and formants are moved separately. Pitch changes by re-spacing the glottal-pulse grains of your voice; each grain's contents are counter-resampled so your vocal-tract resonances stay where they were. The Formant knob exposes that scale directly. See singing and the formants concept page.
What happens to my consonants and breaths? Nothing. Unvoiced sound isn't resynthesised; it passes through as an exact delayed copy of your input, crossfaded with the corrected voice at every boundary.
Is any of this AI? No. Every algorithm is classical, published DSP (pYIN, ZFF, TD-PSOLA, the phase vocoder), implemented in-house. Nothing is learned from your audio, and nothing leaves your machine: no network access, no telemetry.
If I render the same take twice, do I get the same file? Bit-identically, yes. There is no random number generator and no clock-dependent behaviour in the plugin; even the Ensemble "human" scatter is fixed tables. See under the hood.
Why does latency change with Range and Mode? The shifter must buffer about three cycles of the lowest pitch it promises to handle, so the promise sets the wait: raising the floor from 65 Hz to 160 Hz takes Studio mode from ~46 ms to ~19 ms; Live mode's causal grain geometry roughly halves each of those figures, reaching ~10 ms at Range = High. The host is told the exact figure and compensates. See time.
What's the trade in Live mode? Grains are built almost entirely from audio already received: a quarter-period of lookahead instead of Studio's full look around the moment. That's where the ~10 ms comes from. It is PSOLA-only and designed for dry-mic tracking and performance; Studio mode is the record-quality path.
Why do harmony voices come in slightly after the lead? Entrance safety. A newly activated voice holds its gate closed until fresh audio has filled its whole pipeline, then fades in over ~2 ms. The alternative is splicing stale audio, which clicks. The ~25–60 ms entrance lag is the audible signature of that choice. See the crowd.
Why are there two engines? Because grain-domain and spectral-domain shifting have genuinely different textures, and different material flatters different engines. At Range = Full both run warm and latency-aligned, so the Engine switch is a seamless instant A/B. Alternatives ship in the box.
Does it flatten my vibrato? At slow Retune settings the glide follows your pitch contour, so vibrato and bends largely ride through. At instant Retune, flattening them is the intended effect. The tuner never synthesises vibrato of its own.
What does MIDI mode actually lock to? The committed chord (held notes debounced over 30 ms, latched for 250 ms after release), not raw key states, so finger gaps between chords never drop the correction. Lead-note policy is selectable, and every target is folded into the octave you're singing in. See deciding.
Does bypassing shift my timing? No. Bypass switches to a dry path of identical delay, so timing is constant and the transition never clicks.
What happens when my loop restarts? The decision layer resets (glide, smoothing, chord latch) and the audio buffers are left alone, so loops restart without a portamento sweep. (Currently wired for VST3; AU transport plumbing is on the ledger.)
Can it tune two voices at once? A guitar? No. It is a one-voice instrument, by design. Polyphonic material will be heard as one pitch, and corrected as one.
Does it support other tuning references? Yes. Tune spans 415–466 Hz and rescales the entire grid; snapping behaves identically at any reference.
last updated · Charlie