For over a century, speech and voice science has operated on a fundamental assumption: that the vocal tract functions as an acoustic resonator in which sound waves propagate, reflect, and form standing patterns that shape the radiated speech spectrum.[1,2] From Helmholtz’s resonance theory[3] through Fant’s source-filter model[4] to contemporary computational simulations,[5,6] this acoustic paradigm treats the pressure fluctuations within the vocal tract as propagating sound waves—acoustic phenomena governed by the wave equation and traveling at sound velocity (approximately 340 m/s). This seemingly self-evident assumption has become so deeply embedded that it is rarely questioned: the vocal tract is conceptualized as an “acoustic tube” in which formants arise from acoustic resonances, much like the resonant frequencies of a musical wind instrument.[7]
However, a critical examination reveals a fundamental error in this framework: the pressure fluctuations within the vocal tract during phonation are not acoustic waves. They are aerodynamic pressure fields—what we term “pseudo-sound”[8,9]—that propagate at flow velocities (typically 10-50 m/s, an order of magnitude slower than sound velocity; Shadle,[10] remain spatially confined to boundary layers near tissue surfaces, and do not constitute propagating wave phenomena in the classical sense. The transformation from these aerodynamic pressure fluctuations to true radiating acoustic waves occurs primarily at the vocal tract exit, where the flow encounters the impedance discontinuity with ambient air.[9,11] Within the vocal tract itself, pressure fluctuations result from flow-structure interactions: aerodynamic forces acting on tissue surfaces, vortex formation, turbulent eddies, and tissue mechanical responses.[12,13] Acoustic waves do not constitute a governing physical process within the vocal tract during phonation. Treating these phenomena as “acoustic resonances” fundamentally misidentifies their physical nature.
This misidentification has profound consequences. By assuming acoustic wave propagation where none exists, the classical framework attempts to explain tissue-mediated phenomena through acoustic mechanisms—leading to contradictions, ad hoc corrections, and unexplained observations. Voice changes with mucosal edema cannot be attributed to altered acoustic dimensions because the dimensions remain essentially unchanged.[14,15] Paranasal sinus effects on voice quality contradict acoustic theory because the sinuses are decoupled from the main airway.[16,17] Supraglottic tissue compliance effects resist explanation through simple acoustic impedance models.[18,19] The framework works adequately for predicting the radiated spectrum (the acoustic output), but systematically fails to account for the mechanisms by which that output is generated within living tissue.
More strikingly, several well-documented phonatory phenomena remain unexplained or inadequately explained within the acoustic framework: whistle register,[20,21] electrolarynx speech,[22,23] non-periodic voice qualities,[24,25] and extreme register transitions.[26,27] These are not minor anomalies; they represent systematic failures to account for phonatory mechanisms using acoustic principles alone.
Moreover, apparent acoustic phenomena such as the missing fundamental often reflect measurement artifacts rather than physical reality. Standard pitch detection methods infer F0 from spectral spacing, systematically misidentifying source mechanisms when harmonic assumptions are violated.
Attempts to address these issues within the acoustic paradigm have produced increasingly complex hybrid models: source-tract acoustic interaction,[28,29] nonlinear coupling terms,[30] supraglottic impedance corrections,[31] tissue boundary layer effects.[32] Yet these extensions preserve the fundamental assumption of acoustic wave propagation within the tract. They treat tissue properties as boundary conditions modifying an acoustic phenomenon, rather than recognizing that the core phenomenon may be mechanical and aerodynamic rather than acoustic. The result is a proliferation of special cases and correction factors—a classic symptom of a theory stretched beyond its domain of validity.
This paper proposes an alternative framework: Distributed Tissue Vibration. We argue that voice production arises from mechanically coupled tissue vibrations excited by aerodynamic forcing, not from acoustic resonances within an air column. “Formants” ref lect mechanical tissue resonance frequencies where aerodynamic excitation couples efficiently to tissue natural frequencies. The vocal tract functions as a f low-driven mechanical oscillator system, not an acoustic resonator. Radiated sound results from tissue motion at the tract exit generating acoustic waves in the far field—but the formative process within the tract is mechanical and aerodynamic, not acoustic.
The paper proceeds as follows: Section 2 presents empirical evidence from clinical observations and experimental demonstrations. Section 3 classifies voice production theories. Section 4 articulates the non-acoustic framework’s core principles. Section 5 elaborates physical foundations. Section 6 discusses validation methods and clinical applications. Section 7 addresses limitations and proposes empirical tests.
This framework addresses voice and vowel production mechanisms. We do not claim to explain consonant articulation, speech motor control, or prosody. We distinguish between mechanisms within the vocal tract (mechanical) and radiated output (acoustic). Our goal is to resolve contradictions and unexplained phenomena that arise when tissuemechanical processes are misidentified as acoustic resonances.
The purpose of this work is to articulate a coherent biophysical framework grounded in converging observations, physical constraints, and phenomena that resist satisfactory explanation within purely acoustic models. The experimental and clinical observations serve as motivating evidence and mechanistic anchors for theory construction, not as hypothesis-testing experiments.
EVIDENCE BASIS
The Distributed Tissue Vibration framework is grounded in converging clinical observations, experimental demonstrations, and phenomena that resist acoustic explanation. This section presents the empirical foundation motivating the theoretical framework rather than providing definitive experimental proof. The observations serve as mechanistic anchors demonstrating that tissuemechanical interpretations align with existing evidence while explaining anomalies within the acoustic paradigm.
Observational evidence
Multiple lines of clinical and experimental evidence challenge purely acoustic interpretations of voice production. Voice changes with mucosal edema occur despite unchanged vocal tract dimensions, suggesting formant shifts arise from altered tissue mechanical properties (increased mass) rather than geometric changes.[15] Electrolarynx speech produces intelligible formants despite replacing glottal vibration with external mechanical excitation at a fixed frequency and non-physiological location—if formants were purely acoustic cavity resonances, this should not occur.[22,23]
Whistle register exhibits extreme spectral purity (near-sinusoidal waveform) independent of glottal frequency, minimal vocal fold activity with strong supraglottic tissue vibrations, and frequency characteristics unrelated to vocal tract dimensions—consistent with local f low-structure coupling rather than acoustic filtering.[21] Recent tissue imaging reveals that vocal tract tissues vibrate at formant frequencies, not merely at fundamental frequency, with spatial distribution varying by vowel configuration—supporting distributed mechanical resonances.[33]
Non-periodic voice qualities (vocal fry, rough voice, diplophonia) exhibit complex spectral patterns and subharmonics involving tissue mechanical phenomena that resist straightforward acoustic explanation.[24,25] Extreme register transitions show abruptness, spectral discontinuity, and hysteresis suggesting mechanical oscillation regime changes rather than smooth acoustic adjustments.[26]
Experimental Demonstrations
To directly test tissue-mechanical predictions, we conducted experiments examining tissue vibration patterns and mechanical excitation effects on formant generation.
Methods
Tissue Vibration Recordings: Multi-channel piezoelectric contact microphones recorded vibrations at buccal, submental, and prelaryngeal surfaces simultaneously with acoustic output (dynamic microphone). The PANM system (Praat Assisted Nasalance Meter;[34,35] captured nasal/oral aerodynamic and vibratory components. Wavenumber measurements used ReSpeaker 4-Mic Linear Array.
Vibration Generator Experiments: A mechanical vibration transducer (4 Ω/25 W) driven by amplified tone-complex signals (−6 or −12 dB/octave spectral slope, generated in Praat) applied controlled vibrations to neck and facial sites. Articulatory movements without airflow tested whether external mechanical energy alone generates formant-like structure. The VTM-N20 vocal tract model[36] was used to examine vibrational excitation interactions with vocal tract geometry and pseudo-sound generation.
Experimental observations
Muscle Tone Effects on Resonance (Figure 1): Vibration (100 Hz tone complex, −6 dB/octave) applied to inner wrist surface while flexing/extending fingers shows high-frequency energy (6-12 kHz) concentrating during muscle contraction. This demonstrates that muscle tone changes alone, without airflow, reorganize spectral energy—supporting mechanically mediated resonance.
External Vibration Formant Generation (Figure 2): Electrolarynx-like external vibration (100 Hz tone complex, –12 dB/octave) applied submandibular region during articulation of [ɑuɑuɑ] produces welldefined formants at ~300 Hz ([u]) and ~1100 Hz ([ɑ]). Formants arise from tissue mechanical resonances modulated by articulatory posture, independent of acoustic cavity resonances, confirming that formants reflect tissue vibratory response.
Distributed Vibration Beyond Glottis (Figure 3): Multi-channel recordings of [bɑ] show strong buccal tissue vibration during prevoicing phase (before lip release), demonstrating distributed tissue vibration beyond the larynx. Fundamental frequency (~115 Hz) approximates laryngeal F0, consistent with coordinated tissue vibration throughout vocal tract driven by aerodynamic forcing—not localized laryngeal source activity.
Subharmonic Tissue Oscillation (Figure 4): PANM recordings from a child with hypernasal speech during [kɑkɑkɑ] reveal sustained nasal vibration during complete oral occlusion. The nasal signal fundamental frequency (~150 Hz) is one octave below laryngeal F0—supraglottic tissues exhibit autonomous vibratory modes driven by, but not frequency-locked to, laryngeal forcing. This subharmonic vibration represents mechanically governed oscillation, not acoustic resonance.
Ethics statement
Ethics committee approval was not required for this theoretical study. Recordings in Figures 1-3 were self-produced by the author. Figure 4 was obtained with informed parental consent during a routine assessment. All experimental procedures were self-conducted.
THEORETICAL FRAMEWORK: CLASSIFYING VOICE PRODUCTION THEORIES
Voice production theories are often described as progressive refinements of a common framework. However, closer examination reveals distinct theoretical commitments regarding the nature of pressure fluctuations inside the vocal tract. The decisive question is: Does acoustic wave propagation exist within the vocal tract during phonation? Based on this criterion, existing theories can be systematically grouped into four major classes: linear acoustic, nonlinear acoustic, hybrid, and non-acoustic models (Table 1). This classification reflects fundamentally different ontological assumptions about the physical processes underlying voice production.
Linear acoustic theories
Linear acoustic theories assume that the vocal tract behaves as a passive acoustic resonator supporting linear wave propagation, with formants arising from standing acoustic waves within the air column. This framework originates in 19th-century studies by Helmholtz,[3] culminating in Fant’s source-filter model.[4] These theories achieved remarkable success in speech synthesis,[37] acoustic analysis,[1] and perceptual modeling due to their mathematical tractability. However, their explanatory scope is limited: they predict radiated spectra successfully but cannot account for tissue-mechanical phenomena or biomechanical processes of living phonation (Table 1).
Nonlinear acoustic theories
Nonlinear acoustic theories address the inadequacy of source-tract independence while retaining the assumption that acoustic waves propagate within the vocal tract. Rothenberg[31] demonstrated that supraglottal acoustic impedance affects glottal flow and vocal fold vibration. Titze[19,28,38] developed comprehensive theories of nonlinear source-filter coupling, successfully predicting register transitions and formant-harmonic interactions. Despite increased sophistication, these theories preserve the core assumption: resonance remains fundamentally acoustic, with tissue mechanics interpreted as boundary conditions rather than the primary resonant mechanism (Table 1).
Hybrid aerodynamic-acoustic theories
Hybrid theories acknowledge the central role of aerodynamics and tissue mechanics while maintaining that the vocal tract supports acoustic wave propagation. Aeroacoustic approaches[9,11,32,39,40] model sound generation from unsteady flow phenomena, treating aerodynamic fluctuations as sources that generate acoustic waves. Studies revealing pseudo-sound or non-propagating pressure fields[8] are classified here because they interpreted these phenomena as aeroacoustic sources—mechanisms generating acoustic waves rather than replacing them. Modern vibroacoustic interaction studies[33] and FSI models[41] represent the most comprehensive hybrid approach, demonstrating that tissue mechanical impedance significantly affects formant frequencies. However, these models retain a fundamental commitment: resonance phenomena are considered at least partly acoustic (Table 1).
Non-acoustic theories
Non-acoustic theories constitute a fundamental departure: they assume that acoustic wave propagation does not occur within the vocal tract during phonation. Pressure fluctuations are attributed entirely to aerodynamic and mechanical processes— pseudo-sound confined to boundary layers and tissue surfaces. Teager and Teager[42] provided the seminal articulation, reporting that experimental measurements systematically violate classical acoustic impedance relationships. They proposed that speech production is governed by nonlinear vortex dynamics and momentum waves rather than acoustic resonance, directly challenging the acoustic assumption. However, they did not develop a comprehensive alternative framework.
The Distributed Tissue Vibration framework proposed here systematizes this perspective into a complete theory. We argue that formants are not air-column acoustic resonances but mechanical resonance bands of distributed soft tissues driven by aerodynamic forcing. The vocal tract functions as an extended mechanical oscillator system: tissues vibrate in response to aerodynamic pressure fluctuations at frequencies determined by tissue mechanical properties (stiffness, mass, damping) rather than acoustic cavity dimensions. Acoustic waves arise only at the vocal tract exit, where internal flow-related pressure fluctuations and tissue motions transition into radiated far-field sound through an aerodynamicto-acoustic transformation. This framework explains phenomena resisting acoustic explanation (whistle register, electrolarynx speech, tissue compliance effects) as mechanical and aerodynamic processes that only become acoustic upon radiation (Table 1).
Conceptual implications
This classification clarifies that developments in voice theory involve distinct ontological commitments, not merely incremental refinement. Linear, nonlinear, and hybrid theories differ in complexity but share a fundamental assumption: acoustic wave propagation exists within the vocal tract. This shared commitment defines them as variations within the acoustic paradigm.
Non-acoustic theories reject this commitment, proposing a different physical substrate for resonance. Where acoustic theories see standing waves and resonant cavities, non-acoustic theories see flow-driven tissue vibrations and mechanical resonances. Where acoustic theories treat tissue as boundary conditions, non-acoustic theories place tissue mechanics at the explanatory center. If formants arise from tissue mechanical resonances rather than acoustic cavity modes, then tissue property changes naturally affect voice quality despite unchanged cavity dimensions. The accumulated anomalies of acoustic theory become natural consequences of the non-acoustic framework. The question is no longer whether acoustic models can be more complex, but whether acoustic wave propagation within the vocal tract is real or a theoretical assumption that has outlived its empirical support.
CRITICAL DISTINCTIONS: NON-ACOUSTIC FRAMEWORK
Having classified voice production theories into acoustic and non-acoustic paradigms, we now articulate the defining features of the Distributed Tissue Vibration framework. This section focuses on the framework’s core physical commitments: the nature of pseudo-sound, the centrality of tissue vibration, the mechanism of pseudo-sound-tissue coupling, and the distributed character of vibratory phenomena extending beyond the glottis.
Pseudo-sound and aerodynamic phenomena
The framework distinguishes fundamentally between pressure f luctuations within the vocal tract (pseudo-sound) and radiating acoustic waves outside it. Pseudo-sound refers to f low-bound or structure-bound pressure f luctuations that propagate at characteristic f low or deformation velocities (10-50 m/s), not at sound velocity (≈340 m/s). These f luctuations are spatially confined to boundary layers near tissue surfaces and do not constitute propagating acoustic waves governed by the wave equation.[8,40]
Pseudo-sound arises from aerodynamic phenomena including flow instabilities—both periodic (pulsatile f low, shear-layer instabilities, f lute-like regimes, lock-in regimes, rotational structures) and aperiodic (turbulence, chaotic instabilities, wake effects, vortex breakdown). These pressure fluctuations are hydrodynamic in origin—consequences of f luid momentum transport and unsteady flow patterns— rather than acoustic compressions and rarefactions (Table 2). Critically, pseudo-sound cannot radiate to the far field; it remains a near-field phenomenon bound to flow and tissue interfaces. Only at the vocal tract exit does transformation to radiating acoustic waves occur.[9,11]
This distinction resolves the conceptual confusion noted by Teager and Teager:[42] measurements within the vocal tract systematically violate acoustic impedance relationships because the measured pressures are not acoustic. They are aerodynamic pressure fields—hence “pseudo-sound.” Acoustic frameworks misidentify these fluctuations as sound waves, leading to predictive failures when tissue mechanics or flow dynamics dominate.
Speech sounds differ fundamentally in their production requirements. Voiceless consonants rely predominantly on turbulent airflow, while voiced sounds (vowels, voiced consonants, semivowels) depend fundamentally on pseudo-sound-tissue vibration interaction for amplification and sustainability. The source of voicing is not always the glottis: depending on articulation, the oral cavity may dominate in [b], the pharyngeal region in [g], while voiced [h] is directly glottal. This distributed nature aligns with the framework’s core principle that vibration occurs throughout the vocal tract.
The source of harmonics is tissue vibration driven by three forces: (1) glottal pulses (harmonic), (2) elastodynamic waves (harmonic), and (3) rotational phenomena and pseudo-sound (formant). Harmonics are intrinsic properties of tissue vibratory dynamics, not exceptional phenomena requiring complex explanation.
Tissue vibration as the central phenomenon
In acoustic frameworks, tissue serves as boundary conditions—walls defining acoustic cavities. In the non-acoustic framework, tissue is the primary resonant medium. Voice production is fundamentally a mechanical phenomenon: distributed soft tissues vibrate in response to aerodynamic forcing, and these vibrations determine spectral output. The vocal tract is not an acoustic resonator but a mechanical oscillator system.
Tissue vibration encompasses the entire supraglottic airway: pharyngeal walls, tongue body, soft palate, lateral pharyngeal tissues, and oral mucosa. These structures possess mechanical properties—mass, stiffness, damping—that determine their vibrational response. When excited by pseudo-sound pressure fluctuations, tissues vibrate at frequencies determined by their mechanical resonance characteristics, not by acoustic cavity dimensions. Formants reflect these tissue mechanical resonances: frequencies where aerodynamic excitation couples efficiently to tissue natural frequencies.[33]
This framework explains why tissue property changes (edema, inflammation, tension) alter voice quality even when geometric dimensions remain unchanged. Mucosal edema increases tissue mass, shifting mechanical resonance frequencies downward. Muscle tension increases stiffness, raising resonance frequencies. These are mechanical effects on a mechanical system, not acoustic perturbations. Voice disorders are primarily disorders of tissue mechanics— alterations in the mechanical properties that determine vibrational response to aerodynamic forcing.
Bidirectional pseudo-sound-tissue coupling
The framework’s generative mechanism is bidirectional coupling between pseudo-sound and tissue vibration: pseudo-sound generates tissue vibration, and tissue vibration generates pseudo-sound. This reciprocal interaction forms the fundamental dynamic field underlying voice production. Aerodynamic pressure fluctuations (pseudo-sound) exert forces on tissue surfaces, driving vibrations. Simultaneously, tissue motion modulates airflow, creating additional aerodynamic disturbances—new pseudo-sound. This positive feedback can produce sustained, self-excited oscillations even without glottal pulsation. The system is inherently nonlinear: large-amplitude tissue vibrations create strong flow modulation, which reinforces vibration, establishing limit-cycle oscillations.
This coupling mechanism explains phenomena resisting acoustic interpretation. Whistle register arises when a specific f low-tissue resonance achieves strong coupling, producing near-sinusoidal oscillation at a single frequency unrelated to vocal tract dimensions.[21] Electrolarynx speech produces intelligible formants because external mechanical excitation couples to tissue mechanical resonances through the same mechanism, independent of acoustic cavity modes.[23] The framework predicts that artificially exciting tissue vibrations (e.g., via transcutaneous mechanical stimulation) will produce formant-like spectral structure even without airflow, demonstrating that formants are tissue mechanical phenomena.
Distributed vibration beyond the glottis
Acoustic frameworks center voice production at the glottis: the vocal folds are the source, and the tract is a passive filter. The non-acoustic framework rejects this localization. Vibration is distributed throughout the vocal tract. While glottal pulsation initiates flow unsteadiness, subsequent pseudo-sound-tissue interactions occur along the entire airway. Multiple tissue regions can serve as vibratory foci where coupling is particularly strong.
Pharyngeal constriction sites, velopharyngeal port, tongue dorsum position, and lip configurations all create local f low-structure interaction zones. At these locations, tissue vibrations can dominate spectral output. This distributed character explains why formant structure persists in whispered speech (no glottal pulsation): pseudo-sound from turbulent flow at constrictions excites tissue vibrations that produce formant-like spectral peaks. It explains why different vocal tract configurations for the same acoustic formant frequencies produce perceptibly different voice qualities: different tissue regions are vibrating, creating distinct mechanical coupling patterns. The framework predicts that high-speed imaging of pharyngeal and oral tissues during phonation will reveal spatially distributed vibration patterns at formant frequencies, not merely acoustic pressure antinodes.
Acoustic radiation as boundary transformation
Acoustic waves do not exist within the vocal tract during phonation. They arise only at the tract exit through an aerodynamic-to-acoustic transformation. The coupled pseudo-sound-tissue vibration field inside the tract encounters the impedance discontinuity with ambient air at the lips or nostrils. At this boundary, near-field aerodynamic/mechanical phenomena convert to far-field radiating acoustic waves.[9,11]
This transformation is not merely a change in impedance but a change in physical mechanism. Inside: f low-bound pressure f luctuations and tissue mechanical vibrations. Outside: propagating compressional waves in air governed by the wave equation. The radiated spectrum reflects internal pseudo-sound-tissue dynamics but is not a simple acoustic filtering of a source spectrum. It is a hydrodynamic-to-acoustic conversion determined by flow velocities, tissue motion amplitudes, and exit geometry.
This framework explains why voice radiation characteristics depend strongly on lip/nostril configuration: these boundaries determine transformation efficiency from internal dynamics to external acoustics. It predicts that radiated sound directivity patterns arise not from acoustic diffraction but from the spatial distribution of tissue motion and flow patterns at the exit. Different vowels radiate differently not because of different acoustic cavity modes but because tissue vibration patterns differ, creating different transformation dynamics at the boundary.
These five principles—pseudo-sound as non-radiating pressure f luctuations, tissue vibration as the central phenomenon, bidirectional pseudo-sound-tissue coupling, distributed vibration beyond the glottis, and acoustic radiation as boundary transformation—constitute the core of the Distributed Tissue Vibration framework.
CORE PRINCIPLES OF DISTRIBUTED TISSUE VIBRATION
Having established the defining features of the non-acoustic framework, we now outline its physical foundations. This section presents the core principles governing voice production as a flow-driven, tissuemediated mechanical process. These principles— flow unsteadiness, pseudo-sound generation, tissue mechanical response, bidirectional f low-tissue coupling, spatially distributed vibration, and aerodynamic-to-acoustic transformation—constitute the generative basis from which all observed vocal phenomena arise (Table 3).
Voice production requires f low unsteadiness: steady laminar airflow produces no acoustic output regardless of velocity. Temporal variations in velocity, pressure, or flow direction create dynamic pressure fields through glottal pulsation, flow separation at constrictions, and shear layer instabilities.[13,40] This unsteadiness directly couples to tissue mechanics through bidirectional interaction: flow instabilities drive tissue vibrations, and tissue vibrations amplify flow instabilities, producing self-sustaining oscillations when coupling is sufficiently strong.
Pseudo-sound generation occurs through hydrodynamic mechanisms—vortex dynamics, jet instabilities, and flow-structure interaction— creating pressure fluctuations that propagate at flow velocities (10-50 m/s), not sound velocity, and remain confined to boundary layers.[8,11] These pressure fields cannot radiate to the far field; they remain nearfield phenomena. When pseudo-sound frequencies match tissue mechanical resonances, efficient energy coupling produces strong tissue vibrations that dominate spectral output—the mechanism of formant generation.
Vocal tract tissues possess characteristic mechanical properties—mass, stiffness, damping— that determine natural frequencies at which they vibrate most readily under external forcing.[43] When pseudo-sound excites tissues at or near these frequencies, mechanical resonance occurs: vibration amplitude increases dramatically. This mechanical resonance is fundamentally different from acoustic cavity resonance; it depends on tissue properties, not air column geometry, and is governed by structural dynamics, not the wave equation. Formants reflect these tissue mechanical resonances. Tissue properties vary spatially and modulate actively through muscle tension, enabling vowel articulation through frequency shifts. Voice disorders alter mechanical resonance characteristics independent of geometric changes.[44]
The generative core is bidirectional f low-tissue coupling. Unsteady pressure f luctuations (pseudo-sound) exert surface forces driving tissue vibrations; simultaneously, vibrating tissues alter f low geometry, modulating velocity, pressure, and vorticity, generating additional pseudo-sound. This positive feedback produces self-excited oscillations, explaining whistle register (strong coupling at a single frequency), register transitions (geometric changes shifting coupling efficiency), and distributed coupling throughout the vocal tract.[21]
Tissue vibration is not localized to the glottis but distributed throughout the supraglottic tract. Multiple tissue regions vibrate simultaneously at different frequencies corresponding to local mechanical resonances and pseudo-sound patterns. Each articulatory configuration creates unique distributed vibratory patterns, explaining formant persistence in whispered speech and perceptibly different voice qualities for the same acoustic formant frequencies due to distinct tissue vibration patterns.
Acoustic waves arise only at the tract exit through aerodynamic-to-acoustic transformation. The coupled pseudo-sound-tissue vibration field inside encounters the impedance discontinuity with ambient air at lips/nostrils, converting near-field aerodynamic/mechanical phenomena to far-field radiating acoustic waves.[9,11] This is not merely impedance change but physical mechanism change. Inside: flow-bound pressure fluctuations and tissue vibrations. Outside: propagating compressional waves governed by the wave equation. Voice radiation characteristics depend on lip/nostril configuration determining transformation efficiency; directivity patterns arise from spatial distribution of tissue motion and f low patterns at exit, not acoustic diffraction.
Table 3 provides detailed characterization of these six principles, their mechanisms, and roles in voice production. Together, they provide a mechanistic account grounded in f luid dynamics, structural mechanics, and aeroacoustics rather than acoustic wave propagation.
VALIDATION AND CLINICAL TRANSLATION
Having established the theoretical framework and its empirical foundation, this section outlines testable predictions distinguishing the tissue-mechanical approach from acoustic theories, proposes experimental validation methods, and discusses implications for clinical practice.
Testable predictions
The framework makes specific predictions that can empirically distinguish it from acoustic theories:
Prediction 1: Tissue mechanics predicts formants better than geometry. In vivo elastography measuring pharyngeal/supraglottic tissue stiffness during phonation, correlated with simultaneous acoustic recordings, should show that formant variance is better explained by tissue property variance than geometric variance. Studies correlating tissue mechanical measurements with voice acoustics can test whether formant frequencies depend primarily on tissue mechanics or cavity dimensions.
Prediction 2: Mechanical perturbations alter voice independently of geometry. Localized tissue stiffening (e.g., transcutaneous focused ultrasound creating temporary property changes) or mass loading (e.g., topical application of heavy, inert materials) should shift formant frequencies without altering vocal tract shape. Acoustic frameworks predict no effect; the mechanical framework predicts systematic frequency shifts correlating with altered tissue properties.
Prediction 3: Distributed tissue vibration at formant frequencies. High-speed volumetric imaging (optical coherence tomography, ultrasound, emerging MRI techniques) should reveal spatially distributed vibration patterns at formant frequencies throughout vocal tract tissues. Multiple tissue regions should vibrate simultaneously at different formant frequencies, with amplitudes correlating with spectral peak amplitudes. Spatial distribution should vary with articulatory configuration in ways not predictable from acoustic standing wave patterns.
Prediction 4: Artificial tissue excitation generates formants. Transcutaneous mechanical stimulation of pharyngeal/oral tissues at specific frequencies, in the absence of airflow, should produce formant-like spectral peaks in sound recorded near the mouth. This would directly demonstrate that tissue vibrations alone, without acoustic resonances, create formant structure. These experiments are feasible with current technology.
Proposed experimental methods
Advanced imaging and measurement techniques enable direct testing of framework predictions:
Laser Doppler Vibrometry: Non-contact measurement of tissue surface vibration velocities with high spatial and temporal resolution. Can map vibration patterns across pharyngeal and oral tissues during phonation, revealing distributed vibratory fields and their relationship to spectral output.
Shear Wave Elastography: Non-invasive measurement of tissue viscoelastic properties. Can quantify stiffness, damping, and resonance characteristics of vocal tract soft tissues, testing correlations between tissue mechanical properties and formant frequencies.
MR Elastography (MRE): Enables visualization of tissue vibratory patterns during phonation by imaging shear wave propagation through tissues. Can quantify mechanical properties of laryngeal, pharyngeal, and oral tissues—parameters directly relevant to the tissue-mechanical framework.
Phase-Contrast129 Xe MRI: Direct measurement of airflow velocity fields within vocal tract during speech production. Visualizes aerodynamic forcing patterns, validating predictions regarding spatial distribution of aerodynamic energy transfer to tissues and relationship between airflow dynamics and tissue vibration initiation.
Multi-Channel Contact Microphone Arrays: Systematic recordings from distributed vocal tract sites during various phonatory tasks can characterize spatial vibration patterns, phase relationships between tissue regions, and coupling between aerodynamic events and tissue responses.
Clinical applications
The framework transforms clinical practice by prioritizing tissue mechanical properties over geometric measurements.
Voice Assessment: Traditional assessment focuses on vocal tract shape (imaging, acoustic analysis). The framework suggests direct measurement of tissue mechanical properties—stiffness, compliance, damping—as primary diagnostic indicators. Clinical evaluation should include palpation assessment of pharyngeal/supraglottic tissue tension, hydration status (affecting viscoelasticity), and inflammatory markers affecting tissue mechanics. Emerging technologies like shear wave elastography enable non-invasive tissue stiffness measurement; the framework predicts these correlate more strongly with voice quality than geometric parameters.
Voice Therapy: Approaches shift from acoustic targeting to mechanical optimization. Rather than instructing patients to “resonate in the mask” or “place the voice forward” (acoustic metaphors), therapy focuses on reducing tissue mechanical tension, optimizing hydration, and facilitating efficient flow-tissue coupling. Resonant voice therapy succeeds by establishing optimal tissue mechanical states for efficient vibration, not by achieving acoustic resonance. Biofeedback of tissue mechanical properties (e.g., surface electromyography of pharyngeal muscles, visual feedback of tissue motion via endoscopy) may prove more effective than acoustic feedback.
Surgical Considerations: Procedures altering tissue mechanical properties (injection laryngoplasty, tissue augmentation, scarring) have effects beyond geometric restoration. Surgical outcomes depend critically on maintaining appropriate tissue mechanical properties—stiffness matching between native and augmented tissues, preservation of compliance gradients. Material selection for vocal fold augmentation should prioritize mechanical property matching over purely volumetric restoration. Explicitly considering tissue mechanical effects may improve voice outcomes after phonosurgery.
These clinical applications demonstrate the framework’s practical utility, offering new assessment strategies and therapeutic targets that address mechanisms acoustic frameworks cannot explain.
SCOPE, LIMITATIONS, AND OUTLOOK
Framework scope
The framework addresses voice and vowel production mechanisms—phonatory processes involving sustained airflow, tissue vibration, and spectral shaping. It explains formant generation, voice quality variations, register phenomena, and acoustic consequences of tissue mechanical alterations. It does not claim to explain all speech production phenomena. Consonant articulation, particularly obstruent consonants produced through turbulent airflow without significant tissue vibration, may be adequately described by acoustic and aerodynamic frameworks. Speech motor control, articulatory planning, and prosodic timing involve neural and muscular coordination outside the framework’s scope.
The framework is compatible with existing knowledge in complementary domains. It does not reject acoustic phonetics for perception and signal processing—radiated speech is acoustic waves that listeners perceive through auditory mechanisms. It does not dispute articulatory phonetics regarding articulator movements. Rather, it reinterprets the mechanism by which articulatory configurations produce spectral outputs: not through acoustic cavity resonances but through tissue-mechanical resonances modulated by articulatory postures. The framework is additive, not replacive—it offers an alternative explanation for specific phenomena where acoustic frameworks fail while acknowledging acoustic approaches remain valid where they succeed.
Current limitations
Measurement challenges are significant: current technology cannot easily measure tissue mechanical properties in vivo during phonation throughout the vocal tract. Shear wave elastography shows promise but has limited spatial resolution and accessibility to deep pharyngeal tissues. High-speed imaging captures surface motion but cannot reveal internal mechanical properties or three-dimensional vibratory fields.
Computational modeling of the full framework is prohibitively complex. Simulating coupled aerodynamics, tissue structural dynamics, and distributed mechanical resonances requires computational resources beyond current standard practice. While computational fluid dynamics and finite element structural models exist separately, fully coupled simulations incorporating realistic tissue geometry, material properties, and flow conditions remain computationally intensive.
Theoretical gaps persist. Precise mechanisms of pseudo-sound-tissue coupling at specific anatomical sites remain incompletely characterized. The role of tissue hydration in modulating coupling efficiency requires investigation. Individual variability in tissue mechanical properties and their relationship to voice quality is poorly understood. The framework provides qualitative explanations but lacks quantitative predictive models for clinical application.
Concluding remarks
The Distributed Tissue Vibration framework represents a fundamental reconceptualization of voice production. For over a century, acoustic wave propagation within the vocal tract has been accepted as self-evident. We have challenged this assumption, arguing that pressure fluctuations within the tract are aerodynamic (pseudo-sound), that resonance is fundamentally mechanical (tissue properties), and that acoustic waves arise only upon radiation at tract boundaries. This is not an incremental refinement but a paradigm shift—a change in the basic explanatory framework.
Such shifts are resisted not because evidence is lacking but because existing paradigms are entrenched. Acoustic theory works admirably for speech synthesis, signal processing, and perception—domains where radiated output matters. Its failure lies in explaining production mechanisms in living tissue, where mechanical and aerodynamic phenomena dominate. We do not ask researchers to abandon acoustic approaches where they succeed; we ask them to recognize where those approaches systematically fail and to consider that the failure may stem from misidentifying the underlying physics.
The implications extend beyond theory. If voice quality depends primarily on tissue mechanical properties, clinical practice must evolve. Voice assessment should prioritize tissue mechanics; therapy should target mechanical optimization; surgery should preserve mechanical function. If formants arise from tissue resonances, understanding individual voice quality requires understanding individual tissue properties—opening new research directions in personalized voice medicine.
We propose this framework not as final truth but as a testable alternative warranting serious empirical investigation. The predictions are clear, the experiments feasible, the potential impact significant. Whether the framework ultimately prevails or is refined by future evidence, the questions it raises—What is the physical nature of pressure f luctuations in the vocal tract? Do tissues mechanically resonate at formant frequencies? Where do acoustic waves actually begin?—demand answers. Science advances by questioning assumptions, and the acoustic assumption has gone unquestioned long enough.
Finally, we emphasize that radiated acoustics describe the perceptual outcome of voice production, but understanding the underlying physical mechanisms requires direct investigation of tissue mechanics and f low–structure interaction, not inference from acoustic output alone.
