The Fundamental Problem: Time vs. Frequency Coupling (समय बनाम आवृत्ति की समस्या)
In basic digital audio, sound is stored as a series of pulse-code modulation (PCM) amplitude samples over time. If you want to raise the pitch by 50%, the simplest method is to play back the samples 1.5 times faster. However, this compresses the time domain: your speech speeds up like a fast-forwarded cassette tape.
To build a modern real-time voice changer, audio engineers must separate pitch transposition from time duration. This is accomplished through the Phase Vocoder algorithm.
💡 हिंदी सारांश: यदि आप किसी ऑडियो को सिर्फ तेज चलाएंगे तो आवाज पतली तो होगी लेकिन बहुत तेज (फास्ट-फॉरवर्ड) हो जाएगी। आधुनिक वॉइस चेंजर बोलने की गति (स्पीड) को बदले बिना केवल आवाज की पिच और टोन को बदलते हैं।
Step 1: Short-Time Fourier Transform (एसटीएफटी द्वारा फ्रीक्वेंसी विश्लेषण)
The continuous audio stream is chopped into small overlapping frames (typically 1024 or 2048 samples). To avoid sharp spectral edges at frame boundaries, a smoothing envelope such as a Hanning Window is multiplied across the buffer:
Next, the Fast Fourier Transform (FFT) converts this windowed time signal into complex frequency bins, providing two critical arrays:
- Magnitude Spectrum: The energy present at each frequency bin. (प्रत्येक आवृत्ति में ऊर्जा की मात्रा).
- Phase Spectrum: The rotational angle of each frequency bin at that millisecond. (फेज कोण).
Step 2: Phase Unwrapping & Frequency Estimation (फेज अनरैपिंग)
Because the discrete FFT only provides frequency resolution in coarse bins (for instance, every 43Hz at 44.1kHz), true vocal pitches fall between bins. The phase vocoder calculates the phase difference between consecutive frames. By "unwrapping" this phase delta, the algorithm measures the exact instantaneous frequency of each vocal harmonic down to fractions of a Hertz.
Step 3: Resynthesis via Inverse FFT (इन्वर्स एफएफटी और रीसिंथेसिस)
To transpose the voice, the estimated frequencies are multiplied by the desired semitone shift:
New continuous phase angles are accumulated to guarantee that phase alignment remains seamless across frame boundaries. Finally, an Inverse Fast Fourier Transform (IFFT) converts the modified frequency bins back into audible PCM sound waves with zero tempo distortion.
Step 4: Formant Preservation (फॉर्मैंट सुरक्षा और वोकल ट्रैक्ट मॉडलिंग)
Pitch shifting alone makes adult voices sound unnatural. To retain human characteristics, VoiceChanger.in decouples the vocal spectral envelope from the pitch bins. When you choose Deep Voice, we lower the fundamental harmonics while adding an acoustic low-shelf resonance filter at 180Hz to simulate a larger physical vocal tract.
गले और मुंह की गूंज (फॉर्मैंट) को सुरक्षित रखकर ही डीप या गर्ल वॉयस प्राकृतिक लगती है, अन्यथा वह कार्टून जैसी बन जाती है।