Praxis of Otorhinolaryngology

Mehmet Akif Kılıç

Voice-Timbre-Speech-Language Unit, Lokman Hekim İstanbul Hospital, İstanbul, Türkiye

Keywords: Aeroacoustics, flow-structure interaction, phonation, speech biophysics, voice production.

Abstract

For over a century, speech science has assumed that the vocal tract functions as an acoustic resonator, with formants arising from standing waves in the air column. This paper challenges this assumption by demonstrating that pressure fluctuations within the vocal tract during phonation are not acoustic waves but aerodynamic pressure fields (pseudo-sound) that propagate at flow velocities and remain confined to tissue boundary layers. We introduce the Distributed Tissue Vibration (DTV) framework, which relocates resonance from air to tissue. In this view, formants are not acoustic cavity modes but mechanical resonance frequencies of distributed tissues driven by aerodynamic forcing through flow-structure interaction. The transformation into radiating acoustic waves occurs primarily at the vocal tract exit. This framework accounts for phenomena that resist acoustic interpretation, including the whistle register, electrolarynx speech, and voice changes from alterations in tissue properties independent of geometry. Clinical observations, contact microphone recordings, and experiments using a vibration transducer support these tissue-mechanical interpretations. By shifting focus from isolated vocal fold geometry to distributed tissue mechanics across the vocal tract, the DTV model offers a holistic basis for therapy and surgery. It articulates empirical predictions and experimental tests, marking a proposed paradigm shift in voice production theory.

For over a century, speech and voice science has operated on a fundamental assumption: that the vocal tract functions as an acoustic resonator in which sound waves propagate, reflect, and form standing patterns that shape the radiated speech spectrum.[1,2] From Helmholtz’s resonance theory[3] through Fant’s source-filter model[4] to contemporary computational simulations,[5,6] this acoustic paradigm treats the pressure fluctuations within the vocal tract as propagating sound waves—acoustic phenomena governed by the wave equation and traveling at sound velocity (approximately 340 m/s). This seemingly self-evident assumption has become so deeply embedded that it is rarely questioned: the vocal tract is conceptualized as an “acoustic tube” in which formants arise from acoustic resonances, much like the resonant frequencies of a musical wind instrument.[7]

However, a critical examination reveals a fundamental error in this framework: the pressure fluctuations within the vocal tract during phonation are not acoustic waves. They are aerodynamic pressure fields—what we term “pseudo-sound”[8,9]—that propagate at flow velocities (typically 10-50 m/s, an order of magnitude slower than sound velocity; Shadle,[10] remain spatially confined to boundary layers near tissue surfaces, and do not constitute propagating wave phenomena in the classical sense. The transformation from these aerodynamic pressure fluctuations to true radiating acoustic waves occurs primarily at the vocal tract exit, where the flow encounters the impedance discontinuity with ambient air.[9,11] Within the vocal tract itself, pressure fluctuations result from flow-structure interactions: aerodynamic forces acting on tissue surfaces, vortex formation, turbulent eddies, and tissue mechanical responses.[12,13] Acoustic waves do not constitute a governing physical process within the vocal tract during phonation. Treating these phenomena as “acoustic resonances” fundamentally misidentifies their physical nature.

This misidentification has profound consequences. By assuming acoustic wave propagation where none exists, the classical framework attempts to explain tissue-mediated phenomena through acoustic mechanisms—leading to contradictions, ad hoc corrections, and unexplained observations. Voice changes with mucosal edema cannot be attributed to altered acoustic dimensions because the dimensions remain essentially unchanged.[14,15] Paranasal sinus effects on voice quality contradict acoustic theory because the sinuses are decoupled from the main airway.[16,17] Supraglottic tissue compliance effects resist explanation through simple acoustic impedance models.[18,19] The framework works adequately for predicting the radiated spectrum (the acoustic output), but systematically fails to account for the mechanisms by which that output is generated within living tissue.

More strikingly, several well-documented phonatory phenomena remain unexplained or inadequately explained within the acoustic framework: whistle register,[20,21] electrolarynx speech,[22,23] non-periodic voice qualities,[24,25] and extreme register transitions.[26,27] These are not minor anomalies; they represent systematic failures to account for phonatory mechanisms using acoustic principles alone.

Moreover, apparent acoustic phenomena such as the missing fundamental often reflect measurement artifacts rather than physical reality. Standard pitch detection methods infer F0 from spectral spacing, systematically misidentifying source mechanisms when harmonic assumptions are violated.

Attempts to address these issues within the acoustic paradigm have produced increasingly complex hybrid models: source-tract acoustic interaction,[28,29] nonlinear coupling terms,[30] supraglottic impedance corrections,[31] tissue boundary layer effects.[32] Yet these extensions preserve the fundamental assumption of acoustic wave propagation within the tract. They treat tissue properties as boundary conditions modifying an acoustic phenomenon, rather than recognizing that the core phenomenon may be mechanical and aerodynamic rather than acoustic. The result is a proliferation of special cases and correction factors—a classic symptom of a theory stretched beyond its domain of validity.

This paper proposes an alternative framework: Distributed Tissue Vibration. We argue that voice production arises from mechanically coupled tissue vibrations excited by aerodynamic forcing, not from acoustic resonances within an air column. “Formants” ref lect mechanical tissue resonance frequencies where aerodynamic excitation couples efficiently to tissue natural frequencies. The vocal tract functions as a f low-driven mechanical oscillator system, not an acoustic resonator. Radiated sound results from tissue motion at the tract exit generating acoustic waves in the far field—but the formative process within the tract is mechanical and aerodynamic, not acoustic.

The paper proceeds as follows: Section 2 presents empirical evidence from clinical observations and experimental demonstrations. Section 3 classifies voice production theories. Section 4 articulates the non-acoustic framework’s core principles. Section 5 elaborates physical foundations. Section 6 discusses validation methods and clinical applications. Section 7 addresses limitations and proposes empirical tests.

This framework addresses voice and vowel production mechanisms. We do not claim to explain consonant articulation, speech motor control, or prosody. We distinguish between mechanisms within the vocal tract (mechanical) and radiated output (acoustic). Our goal is to resolve contradictions and unexplained phenomena that arise when tissuemechanical processes are misidentified as acoustic resonances.

The purpose of this work is to articulate a coherent biophysical framework grounded in converging observations, physical constraints, and phenomena that resist satisfactory explanation within purely acoustic models. The experimental and clinical observations serve as motivating evidence and mechanistic anchors for theory construction, not as hypothesis-testing experiments.

EVIDENCE BASIS

The Distributed Tissue Vibration framework is grounded in converging clinical observations, experimental demonstrations, and phenomena that resist acoustic explanation. This section presents the empirical foundation motivating the theoretical framework rather than providing definitive experimental proof. The observations serve as mechanistic anchors demonstrating that tissuemechanical interpretations align with existing evidence while explaining anomalies within the acoustic paradigm.

Observational evidence

Multiple lines of clinical and experimental evidence challenge purely acoustic interpretations of voice production. Voice changes with mucosal edema occur despite unchanged vocal tract dimensions, suggesting formant shifts arise from altered tissue mechanical properties (increased mass) rather than geometric changes.[15] Electrolarynx speech produces intelligible formants despite replacing glottal vibration with external mechanical excitation at a fixed frequency and non-physiological location—if formants were purely acoustic cavity resonances, this should not occur.[22,23]

Whistle register exhibits extreme spectral purity (near-sinusoidal waveform) independent of glottal frequency, minimal vocal fold activity with strong supraglottic tissue vibrations, and frequency characteristics unrelated to vocal tract dimensions—consistent with local f low-structure coupling rather than acoustic filtering.[21] Recent tissue imaging reveals that vocal tract tissues vibrate at formant frequencies, not merely at fundamental frequency, with spatial distribution varying by vowel configuration—supporting distributed mechanical resonances.[33]

Non-periodic voice qualities (vocal fry, rough voice, diplophonia) exhibit complex spectral patterns and subharmonics involving tissue mechanical phenomena that resist straightforward acoustic explanation.[24,25] Extreme register transitions show abruptness, spectral discontinuity, and hysteresis suggesting mechanical oscillation regime changes rather than smooth acoustic adjustments.[26]

Experimental Demonstrations

To directly test tissue-mechanical predictions, we conducted experiments examining tissue vibration patterns and mechanical excitation effects on formant generation.

Methods

Tissue Vibration Recordings: Multi-channel piezoelectric contact microphones recorded vibrations at buccal, submental, and prelaryngeal surfaces simultaneously with acoustic output (dynamic microphone). The PANM system (Praat Assisted Nasalance Meter;[34,35] captured nasal/oral aerodynamic and vibratory components. Wavenumber measurements used ReSpeaker 4-Mic Linear Array.

Vibration Generator Experiments: A mechanical vibration transducer (4 Ω/25 W) driven by amplified tone-complex signals (−6 or −12 dB/octave spectral slope, generated in Praat) applied controlled vibrations to neck and facial sites. Articulatory movements without airflow tested whether external mechanical energy alone generates formant-like structure. The VTM-N20 vocal tract model[36] was used to examine vibrational excitation interactions with vocal tract geometry and pseudo-sound generation.

Experimental observations

Muscle Tone Effects on Resonance (Figure 1): Vibration (100 Hz tone complex, −6 dB/octave) applied to inner wrist surface while flexing/extending fingers shows high-frequency energy (6-12 kHz) concentrating during muscle contraction. This demonstrates that muscle tone changes alone, without airflow, reorganize spectral energy—supporting mechanically mediated resonance.

External Vibration Formant Generation (Figure 2): Electrolarynx-like external vibration (100 Hz tone complex, –12 dB/octave) applied submandibular region during articulation of [ɑuɑuɑ] produces welldefined formants at ~300 Hz ([u]) and ~1100 Hz ([ɑ]). Formants arise from tissue mechanical resonances modulated by articulatory posture, independent of acoustic cavity resonances, confirming that formants reflect tissue vibratory response.

Distributed Vibration Beyond Glottis (Figure 3): Multi-channel recordings of [bɑ] show strong buccal tissue vibration during prevoicing phase (before lip release), demonstrating distributed tissue vibration beyond the larynx. Fundamental frequency (~115 Hz) approximates laryngeal F0, consistent with coordinated tissue vibration throughout vocal tract driven by aerodynamic forcing—not localized laryngeal source activity.

Figure 3. Waveform, spectrogram, and fundamental frequency contour of [bɑ] vocalization produced by the author, recorded using a 4-channel system. Channel 1 captures the signal via a dynamic microphone positioned 5 cm from the mouth. Channels 2-4 record tissue vibrations using piezoelectric contact microphones placed on: (2) left buccal skin surface, (3) submental region, and (4) prelaryngeal region. The prevoicing phase preceding the [b] release shows strong tissue vibration signals in the buccal channel, demonstrating that phonatory activity involves distributed tissue vibration beyond the larynx—oral cavity tissues vibrate and contribute to sound production prior to lip release. “onset” marks voicing initiation; “release” marks lip opening. The fundamental frequency during prevoicing approximates the laryngeal F0 (~115 Hz), consistent with coordinated tissue vibration throughout the vocal tract driven by aerodynamic forcing. This observation supports the framework’s principle that voice production involves spatially distributed tissue-mechanical vibrations, rather than localized laryngeal source activity.

Subharmonic Tissue Oscillation (Figure 4): PANM recordings from a child with hypernasal speech during [kɑkɑkɑ] reveal sustained nasal vibration during complete oral occlusion. The nasal signal fundamental frequency (~150 Hz) is one octave below laryngeal F0—supraglottic tissues exhibit autonomous vibratory modes driven by, but not frequency-locked to, laryngeal forcing. This subharmonic vibration represents mechanically governed oscillation, not acoustic resonance.

Ethics statement

Ethics committee approval was not required for this theoretical study. Recordings in Figures 1-3 were self-produced by the author. Figure 4 was obtained with informed parental consent during a routine assessment. All experimental procedures were self-conducted.

THEORETICAL FRAMEWORK: CLASSIFYING VOICE PRODUCTION THEORIES

Voice production theories are often described as progressive refinements of a common framework. However, closer examination reveals distinct theoretical commitments regarding the nature of pressure fluctuations inside the vocal tract. The decisive question is: Does acoustic wave propagation exist within the vocal tract during phonation? Based on this criterion, existing theories can be systematically grouped into four major classes: linear acoustic, nonlinear acoustic, hybrid, and non-acoustic models (Table 1). This classification reflects fundamentally different ontological assumptions about the physical processes underlying voice production.

Linear acoustic theories

Linear acoustic theories assume that the vocal tract behaves as a passive acoustic resonator supporting linear wave propagation, with formants arising from standing acoustic waves within the air column. This framework originates in 19th-century studies by Helmholtz,[3] culminating in Fant’s source-filter model.[4] These theories achieved remarkable success in speech synthesis,[37] acoustic analysis,[1] and perceptual modeling due to their mathematical tractability. However, their explanatory scope is limited: they predict radiated spectra successfully but cannot account for tissue-mechanical phenomena or biomechanical processes of living phonation (Table 1).

Nonlinear acoustic theories

Nonlinear acoustic theories address the inadequacy of source-tract independence while retaining the assumption that acoustic waves propagate within the vocal tract. Rothenberg[31] demonstrated that supraglottal acoustic impedance affects glottal flow and vocal fold vibration. Titze[19,28,38] developed comprehensive theories of nonlinear source-filter coupling, successfully predicting register transitions and formant-harmonic interactions. Despite increased sophistication, these theories preserve the core assumption: resonance remains fundamentally acoustic, with tissue mechanics interpreted as boundary conditions rather than the primary resonant mechanism (Table 1).

Hybrid aerodynamic-acoustic theories

Hybrid theories acknowledge the central role of aerodynamics and tissue mechanics while maintaining that the vocal tract supports acoustic wave propagation. Aeroacoustic approaches[9,11,32,39,40] model sound generation from unsteady flow phenomena, treating aerodynamic fluctuations as sources that generate acoustic waves. Studies revealing pseudo-sound or non-propagating pressure fields[8] are classified here because they interpreted these phenomena as aeroacoustic sources—mechanisms generating acoustic waves rather than replacing them. Modern vibroacoustic interaction studies[33] and FSI models[41] represent the most comprehensive hybrid approach, demonstrating that tissue mechanical impedance significantly affects formant frequencies. However, these models retain a fundamental commitment: resonance phenomena are considered at least partly acoustic (Table 1).

Non-acoustic theories

Non-acoustic theories constitute a fundamental departure: they assume that acoustic wave propagation does not occur within the vocal tract during phonation. Pressure fluctuations are attributed entirely to aerodynamic and mechanical processes— pseudo-sound confined to boundary layers and tissue surfaces. Teager and Teager[42] provided the seminal articulation, reporting that experimental measurements systematically violate classical acoustic impedance relationships. They proposed that speech production is governed by nonlinear vortex dynamics and momentum waves rather than acoustic resonance, directly challenging the acoustic assumption. However, they did not develop a comprehensive alternative framework.

The Distributed Tissue Vibration framework proposed here systematizes this perspective into a complete theory. We argue that formants are not air-column acoustic resonances but mechanical resonance bands of distributed soft tissues driven by aerodynamic forcing. The vocal tract functions as an extended mechanical oscillator system: tissues vibrate in response to aerodynamic pressure fluctuations at frequencies determined by tissue mechanical properties (stiffness, mass, damping) rather than acoustic cavity dimensions. Acoustic waves arise only at the vocal tract exit, where internal flow-related pressure fluctuations and tissue motions transition into radiated far-field sound through an aerodynamicto-acoustic transformation. This framework explains phenomena resisting acoustic explanation (whistle register, electrolarynx speech, tissue compliance effects) as mechanical and aerodynamic processes that only become acoustic upon radiation (Table 1).

Conceptual implications

This classification clarifies that developments in voice theory involve distinct ontological commitments, not merely incremental refinement. Linear, nonlinear, and hybrid theories differ in complexity but share a fundamental assumption: acoustic wave propagation exists within the vocal tract. This shared commitment defines them as variations within the acoustic paradigm.

Non-acoustic theories reject this commitment, proposing a different physical substrate for resonance. Where acoustic theories see standing waves and resonant cavities, non-acoustic theories see flow-driven tissue vibrations and mechanical resonances. Where acoustic theories treat tissue as boundary conditions, non-acoustic theories place tissue mechanics at the explanatory center. If formants arise from tissue mechanical resonances rather than acoustic cavity modes, then tissue property changes naturally affect voice quality despite unchanged cavity dimensions. The accumulated anomalies of acoustic theory become natural consequences of the non-acoustic framework. The question is no longer whether acoustic models can be more complex, but whether acoustic wave propagation within the vocal tract is real or a theoretical assumption that has outlived its empirical support.

CRITICAL DISTINCTIONS: NON-ACOUSTIC FRAMEWORK

Having classified voice production theories into acoustic and non-acoustic paradigms, we now articulate the defining features of the Distributed Tissue Vibration framework. This section focuses on the framework’s core physical commitments: the nature of pseudo-sound, the centrality of tissue vibration, the mechanism of pseudo-sound-tissue coupling, and the distributed character of vibratory phenomena extending beyond the glottis.

Pseudo-sound and aerodynamic phenomena

The framework distinguishes fundamentally between pressure f luctuations within the vocal tract (pseudo-sound) and radiating acoustic waves outside it. Pseudo-sound refers to f low-bound or structure-bound pressure f luctuations that propagate at characteristic f low or deformation velocities (10-50 m/s), not at sound velocity (≈340 m/s). These f luctuations are spatially confined to boundary layers near tissue surfaces and do not constitute propagating acoustic waves governed by the wave equation.[8,40]

Pseudo-sound arises from aerodynamic phenomena including flow instabilities—both periodic (pulsatile f low, shear-layer instabilities, f lute-like regimes, lock-in regimes, rotational structures) and aperiodic (turbulence, chaotic instabilities, wake effects, vortex breakdown). These pressure fluctuations are hydrodynamic in origin—consequences of f luid momentum transport and unsteady flow patterns— rather than acoustic compressions and rarefactions (Table 2). Critically, pseudo-sound cannot radiate to the far field; it remains a near-field phenomenon bound to flow and tissue interfaces. Only at the vocal tract exit does transformation to radiating acoustic waves occur.[9,11]

This distinction resolves the conceptual confusion noted by Teager and Teager:[42] measurements within the vocal tract systematically violate acoustic impedance relationships because the measured pressures are not acoustic. They are aerodynamic pressure fields—hence “pseudo-sound.” Acoustic frameworks misidentify these fluctuations as sound waves, leading to predictive failures when tissue mechanics or flow dynamics dominate.

Speech sounds differ fundamentally in their production requirements. Voiceless consonants rely predominantly on turbulent airflow, while voiced sounds (vowels, voiced consonants, semivowels) depend fundamentally on pseudo-sound-tissue vibration interaction for amplification and sustainability. The source of voicing is not always the glottis: depending on articulation, the oral cavity may dominate in [b], the pharyngeal region in [g], while voiced [h] is directly glottal. This distributed nature aligns with the framework’s core principle that vibration occurs throughout the vocal tract.

The source of harmonics is tissue vibration driven by three forces: (1) glottal pulses (harmonic), (2) elastodynamic waves (harmonic), and (3) rotational phenomena and pseudo-sound (formant). Harmonics are intrinsic properties of tissue vibratory dynamics, not exceptional phenomena requiring complex explanation.

Tissue vibration as the central phenomenon

In acoustic frameworks, tissue serves as boundary conditions—walls defining acoustic cavities. In the non-acoustic framework, tissue is the primary resonant medium. Voice production is fundamentally a mechanical phenomenon: distributed soft tissues vibrate in response to aerodynamic forcing, and these vibrations determine spectral output. The vocal tract is not an acoustic resonator but a mechanical oscillator system.

Tissue vibration encompasses the entire supraglottic airway: pharyngeal walls, tongue body, soft palate, lateral pharyngeal tissues, and oral mucosa. These structures possess mechanical properties—mass, stiffness, damping—that determine their vibrational response. When excited by pseudo-sound pressure fluctuations, tissues vibrate at frequencies determined by their mechanical resonance characteristics, not by acoustic cavity dimensions. Formants reflect these tissue mechanical resonances: frequencies where aerodynamic excitation couples efficiently to tissue natural frequencies.[33]

This framework explains why tissue property changes (edema, inflammation, tension) alter voice quality even when geometric dimensions remain unchanged. Mucosal edema increases tissue mass, shifting mechanical resonance frequencies downward. Muscle tension increases stiffness, raising resonance frequencies. These are mechanical effects on a mechanical system, not acoustic perturbations. Voice disorders are primarily disorders of tissue mechanics— alterations in the mechanical properties that determine vibrational response to aerodynamic forcing.

Bidirectional pseudo-sound-tissue coupling

The framework’s generative mechanism is bidirectional coupling between pseudo-sound and tissue vibration: pseudo-sound generates tissue vibration, and tissue vibration generates pseudo-sound. This reciprocal interaction forms the fundamental dynamic field underlying voice production. Aerodynamic pressure fluctuations (pseudo-sound) exert forces on tissue surfaces, driving vibrations. Simultaneously, tissue motion modulates airflow, creating additional aerodynamic disturbances—new pseudo-sound. This positive feedback can produce sustained, self-excited oscillations even without glottal pulsation. The system is inherently nonlinear: large-amplitude tissue vibrations create strong flow modulation, which reinforces vibration, establishing limit-cycle oscillations.

This coupling mechanism explains phenomena resisting acoustic interpretation. Whistle register arises when a specific f low-tissue resonance achieves strong coupling, producing near-sinusoidal oscillation at a single frequency unrelated to vocal tract dimensions.[21] Electrolarynx speech produces intelligible formants because external mechanical excitation couples to tissue mechanical resonances through the same mechanism, independent of acoustic cavity modes.[23] The framework predicts that artificially exciting tissue vibrations (e.g., via transcutaneous mechanical stimulation) will produce formant-like spectral structure even without airflow, demonstrating that formants are tissue mechanical phenomena.

Distributed vibration beyond the glottis

Acoustic frameworks center voice production at the glottis: the vocal folds are the source, and the tract is a passive filter. The non-acoustic framework rejects this localization. Vibration is distributed throughout the vocal tract. While glottal pulsation initiates flow unsteadiness, subsequent pseudo-sound-tissue interactions occur along the entire airway. Multiple tissue regions can serve as vibratory foci where coupling is particularly strong.

Pharyngeal constriction sites, velopharyngeal port, tongue dorsum position, and lip configurations all create local f low-structure interaction zones. At these locations, tissue vibrations can dominate spectral output. This distributed character explains why formant structure persists in whispered speech (no glottal pulsation): pseudo-sound from turbulent flow at constrictions excites tissue vibrations that produce formant-like spectral peaks. It explains why different vocal tract configurations for the same acoustic formant frequencies produce perceptibly different voice qualities: different tissue regions are vibrating, creating distinct mechanical coupling patterns. The framework predicts that high-speed imaging of pharyngeal and oral tissues during phonation will reveal spatially distributed vibration patterns at formant frequencies, not merely acoustic pressure antinodes.

Acoustic radiation as boundary transformation

Acoustic waves do not exist within the vocal tract during phonation. They arise only at the tract exit through an aerodynamic-to-acoustic transformation. The coupled pseudo-sound-tissue vibration field inside the tract encounters the impedance discontinuity with ambient air at the lips or nostrils. At this boundary, near-field aerodynamic/mechanical phenomena convert to far-field radiating acoustic waves.[9,11]

This transformation is not merely a change in impedance but a change in physical mechanism. Inside: f low-bound pressure f luctuations and tissue mechanical vibrations. Outside: propagating compressional waves in air governed by the wave equation. The radiated spectrum reflects internal pseudo-sound-tissue dynamics but is not a simple acoustic filtering of a source spectrum. It is a hydrodynamic-to-acoustic conversion determined by flow velocities, tissue motion amplitudes, and exit geometry.

This framework explains why voice radiation characteristics depend strongly on lip/nostril configuration: these boundaries determine transformation efficiency from internal dynamics to external acoustics. It predicts that radiated sound directivity patterns arise not from acoustic diffraction but from the spatial distribution of tissue motion and flow patterns at the exit. Different vowels radiate differently not because of different acoustic cavity modes but because tissue vibration patterns differ, creating different transformation dynamics at the boundary.

These five principles—pseudo-sound as non-radiating pressure f luctuations, tissue vibration as the central phenomenon, bidirectional pseudo-sound-tissue coupling, distributed vibration beyond the glottis, and acoustic radiation as boundary transformation—constitute the core of the Distributed Tissue Vibration framework.

CORE PRINCIPLES OF DISTRIBUTED TISSUE VIBRATION

Having established the defining features of the non-acoustic framework, we now outline its physical foundations. This section presents the core principles governing voice production as a flow-driven, tissuemediated mechanical process. These principles— flow unsteadiness, pseudo-sound generation, tissue mechanical response, bidirectional f low-tissue coupling, spatially distributed vibration, and aerodynamic-to-acoustic transformation—constitute the generative basis from which all observed vocal phenomena arise (Table 3).

Voice production requires f low unsteadiness: steady laminar airflow produces no acoustic output regardless of velocity. Temporal variations in velocity, pressure, or flow direction create dynamic pressure fields through glottal pulsation, flow separation at constrictions, and shear layer instabilities.[13,40] This unsteadiness directly couples to tissue mechanics through bidirectional interaction: flow instabilities drive tissue vibrations, and tissue vibrations amplify flow instabilities, producing self-sustaining oscillations when coupling is sufficiently strong.

Pseudo-sound generation occurs through hydrodynamic mechanisms—vortex dynamics, jet instabilities, and flow-structure interaction— creating pressure fluctuations that propagate at flow velocities (10-50 m/s), not sound velocity, and remain confined to boundary layers.[8,11] These pressure fields cannot radiate to the far field; they remain nearfield phenomena. When pseudo-sound frequencies match tissue mechanical resonances, efficient energy coupling produces strong tissue vibrations that dominate spectral output—the mechanism of formant generation.

Vocal tract tissues possess characteristic mechanical properties—mass, stiffness, damping— that determine natural frequencies at which they vibrate most readily under external forcing.[43] When pseudo-sound excites tissues at or near these frequencies, mechanical resonance occurs: vibration amplitude increases dramatically. This mechanical resonance is fundamentally different from acoustic cavity resonance; it depends on tissue properties, not air column geometry, and is governed by structural dynamics, not the wave equation. Formants reflect these tissue mechanical resonances. Tissue properties vary spatially and modulate actively through muscle tension, enabling vowel articulation through frequency shifts. Voice disorders alter mechanical resonance characteristics independent of geometric changes.[44]

The generative core is bidirectional f low-tissue coupling. Unsteady pressure f luctuations (pseudo-sound) exert surface forces driving tissue vibrations; simultaneously, vibrating tissues alter f low geometry, modulating velocity, pressure, and vorticity, generating additional pseudo-sound. This positive feedback produces self-excited oscillations, explaining whistle register (strong coupling at a single frequency), register transitions (geometric changes shifting coupling efficiency), and distributed coupling throughout the vocal tract.[21]

Tissue vibration is not localized to the glottis but distributed throughout the supraglottic tract. Multiple tissue regions vibrate simultaneously at different frequencies corresponding to local mechanical resonances and pseudo-sound patterns. Each articulatory configuration creates unique distributed vibratory patterns, explaining formant persistence in whispered speech and perceptibly different voice qualities for the same acoustic formant frequencies due to distinct tissue vibration patterns.

Acoustic waves arise only at the tract exit through aerodynamic-to-acoustic transformation. The coupled pseudo-sound-tissue vibration field inside encounters the impedance discontinuity with ambient air at lips/nostrils, converting near-field aerodynamic/mechanical phenomena to far-field radiating acoustic waves.[9,11] This is not merely impedance change but physical mechanism change. Inside: flow-bound pressure fluctuations and tissue vibrations. Outside: propagating compressional waves governed by the wave equation. Voice radiation characteristics depend on lip/nostril configuration determining transformation efficiency; directivity patterns arise from spatial distribution of tissue motion and f low patterns at exit, not acoustic diffraction.

Table 3 provides detailed characterization of these six principles, their mechanisms, and roles in voice production. Together, they provide a mechanistic account grounded in f luid dynamics, structural mechanics, and aeroacoustics rather than acoustic wave propagation.

VALIDATION AND CLINICAL TRANSLATION

Having established the theoretical framework and its empirical foundation, this section outlines testable predictions distinguishing the tissue-mechanical approach from acoustic theories, proposes experimental validation methods, and discusses implications for clinical practice.

Testable predictions

The framework makes specific predictions that can empirically distinguish it from acoustic theories:

Prediction 1: Tissue mechanics predicts formants better than geometry. In vivo elastography measuring pharyngeal/supraglottic tissue stiffness during phonation, correlated with simultaneous acoustic recordings, should show that formant variance is better explained by tissue property variance than geometric variance. Studies correlating tissue mechanical measurements with voice acoustics can test whether formant frequencies depend primarily on tissue mechanics or cavity dimensions.

Prediction 2: Mechanical perturbations alter voice independently of geometry. Localized tissue stiffening (e.g., transcutaneous focused ultrasound creating temporary property changes) or mass loading (e.g., topical application of heavy, inert materials) should shift formant frequencies without altering vocal tract shape. Acoustic frameworks predict no effect; the mechanical framework predicts systematic frequency shifts correlating with altered tissue properties.

Prediction 3: Distributed tissue vibration at formant frequencies. High-speed volumetric imaging (optical coherence tomography, ultrasound, emerging MRI techniques) should reveal spatially distributed vibration patterns at formant frequencies throughout vocal tract tissues. Multiple tissue regions should vibrate simultaneously at different formant frequencies, with amplitudes correlating with spectral peak amplitudes. Spatial distribution should vary with articulatory configuration in ways not predictable from acoustic standing wave patterns.

Prediction 4: Artificial tissue excitation generates formants. Transcutaneous mechanical stimulation of pharyngeal/oral tissues at specific frequencies, in the absence of airflow, should produce formant-like spectral peaks in sound recorded near the mouth. This would directly demonstrate that tissue vibrations alone, without acoustic resonances, create formant structure. These experiments are feasible with current technology.

Proposed experimental methods

Advanced imaging and measurement techniques enable direct testing of framework predictions:

Laser Doppler Vibrometry: Non-contact measurement of tissue surface vibration velocities with high spatial and temporal resolution. Can map vibration patterns across pharyngeal and oral tissues during phonation, revealing distributed vibratory fields and their relationship to spectral output.

Shear Wave Elastography: Non-invasive measurement of tissue viscoelastic properties. Can quantify stiffness, damping, and resonance characteristics of vocal tract soft tissues, testing correlations between tissue mechanical properties and formant frequencies.

MR Elastography (MRE): Enables visualization of tissue vibratory patterns during phonation by imaging shear wave propagation through tissues. Can quantify mechanical properties of laryngeal, pharyngeal, and oral tissues—parameters directly relevant to the tissue-mechanical framework.

Phase-Contrast129 Xe MRI: Direct measurement of airflow velocity fields within vocal tract during speech production. Visualizes aerodynamic forcing patterns, validating predictions regarding spatial distribution of aerodynamic energy transfer to tissues and relationship between airflow dynamics and tissue vibration initiation.

Multi-Channel Contact Microphone Arrays: Systematic recordings from distributed vocal tract sites during various phonatory tasks can characterize spatial vibration patterns, phase relationships between tissue regions, and coupling between aerodynamic events and tissue responses.

Clinical applications

The framework transforms clinical practice by prioritizing tissue mechanical properties over geometric measurements.

Voice Assessment: Traditional assessment focuses on vocal tract shape (imaging, acoustic analysis). The framework suggests direct measurement of tissue mechanical properties—stiffness, compliance, damping—as primary diagnostic indicators. Clinical evaluation should include palpation assessment of pharyngeal/supraglottic tissue tension, hydration status (affecting viscoelasticity), and inflammatory markers affecting tissue mechanics. Emerging technologies like shear wave elastography enable non-invasive tissue stiffness measurement; the framework predicts these correlate more strongly with voice quality than geometric parameters.

Voice Therapy: Approaches shift from acoustic targeting to mechanical optimization. Rather than instructing patients to “resonate in the mask” or “place the voice forward” (acoustic metaphors), therapy focuses on reducing tissue mechanical tension, optimizing hydration, and facilitating efficient flow-tissue coupling. Resonant voice therapy succeeds by establishing optimal tissue mechanical states for efficient vibration, not by achieving acoustic resonance. Biofeedback of tissue mechanical properties (e.g., surface electromyography of pharyngeal muscles, visual feedback of tissue motion via endoscopy) may prove more effective than acoustic feedback.

Surgical Considerations: Procedures altering tissue mechanical properties (injection laryngoplasty, tissue augmentation, scarring) have effects beyond geometric restoration. Surgical outcomes depend critically on maintaining appropriate tissue mechanical properties—stiffness matching between native and augmented tissues, preservation of compliance gradients. Material selection for vocal fold augmentation should prioritize mechanical property matching over purely volumetric restoration. Explicitly considering tissue mechanical effects may improve voice outcomes after phonosurgery.

These clinical applications demonstrate the framework’s practical utility, offering new assessment strategies and therapeutic targets that address mechanisms acoustic frameworks cannot explain.

SCOPE, LIMITATIONS, AND OUTLOOK

Framework scope

The framework addresses voice and vowel production mechanisms—phonatory processes involving sustained airflow, tissue vibration, and spectral shaping. It explains formant generation, voice quality variations, register phenomena, and acoustic consequences of tissue mechanical alterations. It does not claim to explain all speech production phenomena. Consonant articulation, particularly obstruent consonants produced through turbulent airflow without significant tissue vibration, may be adequately described by acoustic and aerodynamic frameworks. Speech motor control, articulatory planning, and prosodic timing involve neural and muscular coordination outside the framework’s scope.

The framework is compatible with existing knowledge in complementary domains. It does not reject acoustic phonetics for perception and signal processing—radiated speech is acoustic waves that listeners perceive through auditory mechanisms. It does not dispute articulatory phonetics regarding articulator movements. Rather, it reinterprets the mechanism by which articulatory configurations produce spectral outputs: not through acoustic cavity resonances but through tissue-mechanical resonances modulated by articulatory postures. The framework is additive, not replacive—it offers an alternative explanation for specific phenomena where acoustic frameworks fail while acknowledging acoustic approaches remain valid where they succeed.

Current limitations

Measurement challenges are significant: current technology cannot easily measure tissue mechanical properties in vivo during phonation throughout the vocal tract. Shear wave elastography shows promise but has limited spatial resolution and accessibility to deep pharyngeal tissues. High-speed imaging captures surface motion but cannot reveal internal mechanical properties or three-dimensional vibratory fields.

Computational modeling of the full framework is prohibitively complex. Simulating coupled aerodynamics, tissue structural dynamics, and distributed mechanical resonances requires computational resources beyond current standard practice. While computational fluid dynamics and finite element structural models exist separately, fully coupled simulations incorporating realistic tissue geometry, material properties, and flow conditions remain computationally intensive.

Theoretical gaps persist. Precise mechanisms of pseudo-sound-tissue coupling at specific anatomical sites remain incompletely characterized. The role of tissue hydration in modulating coupling efficiency requires investigation. Individual variability in tissue mechanical properties and their relationship to voice quality is poorly understood. The framework provides qualitative explanations but lacks quantitative predictive models for clinical application.

Concluding remarks

The Distributed Tissue Vibration framework represents a fundamental reconceptualization of voice production. For over a century, acoustic wave propagation within the vocal tract has been accepted as self-evident. We have challenged this assumption, arguing that pressure fluctuations within the tract are aerodynamic (pseudo-sound), that resonance is fundamentally mechanical (tissue properties), and that acoustic waves arise only upon radiation at tract boundaries. This is not an incremental refinement but a paradigm shift—a change in the basic explanatory framework.

Such shifts are resisted not because evidence is lacking but because existing paradigms are entrenched. Acoustic theory works admirably for speech synthesis, signal processing, and perception—domains where radiated output matters. Its failure lies in explaining production mechanisms in living tissue, where mechanical and aerodynamic phenomena dominate. We do not ask researchers to abandon acoustic approaches where they succeed; we ask them to recognize where those approaches systematically fail and to consider that the failure may stem from misidentifying the underlying physics.

The implications extend beyond theory. If voice quality depends primarily on tissue mechanical properties, clinical practice must evolve. Voice assessment should prioritize tissue mechanics; therapy should target mechanical optimization; surgery should preserve mechanical function. If formants arise from tissue resonances, understanding individual voice quality requires understanding individual tissue properties—opening new research directions in personalized voice medicine.

We propose this framework not as final truth but as a testable alternative warranting serious empirical investigation. The predictions are clear, the experiments feasible, the potential impact significant. Whether the framework ultimately prevails or is refined by future evidence, the questions it raises—What is the physical nature of pressure f luctuations in the vocal tract? Do tissues mechanically resonate at formant frequencies? Where do acoustic waves actually begin?—demand answers. Science advances by questioning assumptions, and the acoustic assumption has gone unquestioned long enough.

Finally, we emphasize that radiated acoustics describe the perceptual outcome of voice production, but understanding the underlying physical mechanisms requires direct investigation of tissue mechanics and f low–structure interaction, not inference from acoustic output alone.

Citation: Kılıç MA. From acoustic resonance to flow-induced distributed tissue vibration: An aero-elasto-dynamic theoretical framework for voice and speech production. Praxis Otorhinolaryngol 2026;14(3):186-202. https://doi.org/10.5606/kbbu.2026.25.

Data Sharing Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.

AI Disclosure
The author declare that artificial intelligence (AI) tools were not used, or were used solely for language editing, and had no role in data analysis, interpretation, or the formulation of conclusions. All scientific content, data interpretation, and conclusions are the sole responsibility of the author. The author further confirm that AI tools were not used to generate, fabricate, or ‘hallucinate’ references, and that all references have been carefully verified for accuracy.

Conflict of Interest

The author declared no conflicts of interest with respect to the authorship and/or publication of this article.

Financial Disclosure

The author received no financial support for the research and/or authorship of this article.

References

  1. Stevens KN. Acoustic phonetics. Cambridge, MA: MIT Press; 2000.
  2. Titze IR. Principles of voice production. Iowa City, IA: National Center for Voice and Speech; 2000.
  3. Helmholtz H. On the sensations of tone as a physiological basis for the theory of music. Translated: Alexander J. Ellis. London: Longmans, Green, and Co.; 1875.
  4. Fant G. Acoustic theory of speech production: with calculations based on X-ray studies of Russian articulations. The Hague: Mouton; 1970.
  5. Story BH. A parametric model of the vocal tract area function for vowel and consonant simulation. J Acoust Soc Am 2005;117:3231-54. doi: 10.1121/1.1869752.
  6. Vampola T, Horáček J, Švec JG. FE modeling of human vocal tract acoustics. Part I: Production of Czech vowels. Acta Acust United Acust 2008;94:433-47. doi: 10.3813/AAA.918051.
  7. Sundberg J. The science of the singing voice. DeKalb, IL: Northern Illinois University Press; 1987.
  8. Barney A, Shadle CH, Davies POAL. Fluid flow in a dynamic mechanical model of the vocal folds and tract. I. Measurements and theory. J Acoust Soc Am 1999;105:444- 55. doi: 10.1121/1.424504.
  9. Howe MS, McGowan RS. Aeroacoustics of [s]. Proc R Soc A Math Phys Eng Sci 2005;461:1005-1028. doi: 10.1098/ rspa.2004.1405.
  10. Shadle CH. The acoustics of fricative consonants [Dissertation]. Cambridge, MA: Massachusetts Institute of Technology; 1985.
  11. Powell A. Theory of vortex sound. J Acoust Soc Am 1964;36:177-95.
  12. Mongeau L, Franchek N, Coker CH, Kubli RA. Characteristics of a pulsating jet through a small modulated orifice, with application to voice production. J Acoust Soc Am 1997;102:1121-33. doi: 10.1121/1.419864.
  13. Zhang C, Zhao W, Frankel SH, Mongeau L. Computational aeroacoustics of phonation, part II: Effects of flow parameters and ventricular folds. J Acoust Soc Am 2002;112:2147-54. doi: 10.1121/1.1506694.
  14. Titze IR, Lemke J, Montequin D. Populations in the U.S. workforce who rely on voice as a primary tool of trade: A preliminary report. J Voice 1997;11:254-9. doi: 10.1016/ s0892-1997(97)80002-1.
  15. Roy N, Merrill RM, Thibeault S, Gray SD, Smith EM. Voice disorders in teachers and the general population: effects on work performance, attendance, and future career choices. J Speech Lang Hear Res 2004;47:542-51. doi: 10.1044/1092-4388(2004/042).
  16. Flanagan JL. Speech analysis synthesis and perception. Berlin, Heidelberg: Springer Berlin Heidelberg; 1972.
  17. Dang J, Honda K. Acoustic characteristics of the human paranasal sinuses derived from transmission characteristic measurement and morphological observation. J Acoust Soc Am 1996;100:3374-83. doi: 10.1121/1.416978.
  18. Story BH, Titze IR, Hoffman EA. Vocal tract area functions from magnetic resonance imaging. J Acoust Soc Am 1996;100:537-54. doi: 10.1121/1.415960.
  19. Titze IR. Nonlinear source-filter coupling in phonation: theory. J Acoust Soc Am 2008;123:2733-49. doi: 10.1121/1.2832337.
  20. Svec JG, Schutte HK, Miller DG. On pitch jumps between chest and falsetto registers in voice: Data from living and excised human larynges. J Acoust Soc Am 1999;106:1523- 31. doi: 10.1121/1.427149.
  21. Miller DG, Schutte HK. Physical definition of the “flageolet register”. J Voice 1993;7:206-12. doi: 10.1016/ s0892-1997(05)80328-5.
  22. Weiss MS, Yeni-Komshian GH, Heinz JM. Acoustical and perceptual characteristics of speech produced with an electronic artificial larynx. J Acoust Soc Am 1979;65:1298- 308. doi: 10.1121/1.382697.
  23. Liu H, Ng ML. Electrolarynx in voice rehabilitation. Auris Nasus Larynx 2007;34:327-32. doi: 10.1016/j. anl.2006.11.010.
  24. Berry DA, Herzel H, Titze IR, Krischer K. Interpretation of biomechanical simulations of normal and chaotic vocal fold oscillations with empirical eigenfunctions. J Acoust Soc Am 1994;95:3595-604. doi: 10.1121/1.409875.
  25. Herzel H, Berry D, Titze IR, Saleh M. Analysis of vocal disorders with methods from nonlinear dynamics. J Speech Hear Res 1994;37:1008-19. doi: 10.1044/jshr.3705.1008.
  26. Roubeau B, Henrich N, Castellengo M. Laryngeal vibratory mechanisms: The notion of vocal register revisited. J Voice 2009;23:425-38. doi: 10.1016/j.jvoice.2007.10.014.
  27. Henrich N, D'Alessandro C, Doval B, Castellengo M. Glottal open quotient in singing: Measurements and correlation with laryngeal mechanisms, vocal intensity, and fundamental frequency. J Acoust Soc Am 2005;117:1417- 30. doi: 10.1121/1.1850031.
  28. Titze IR, Alipour F. The myoelastic aerodynamic theory of phonation. Iowa City, IA: National Center for Voice and Speech; 2006.
  29. Story BH. An overview of the physiology, physics and modeling of the sound source for vowels. Acoust Sci Technol 2002;23:195-206.
  30. Zhang Z, Mongeau L, Frankel SH, Thomson S, Park JB. Sound generation by steady flow through glottisshaped orifices. J Acoust Soc Am 2004;116:1720-8. doi: 10.1121/1.1779331.
  31. Rothenberg M. Acoustic interaction between the glottal source and the vocal tract. In: Stevens KN, Hirano M, editors. Vocal fold physiology. Tokyo: University of Tokyo Press; 1981. p. 305-28.
  32. McGowan RS. An aeroacoustic approach to phonation. J Acoust Soc Am 1988;83:696-704. doi: 10.1121/1.396165.
  33. Fleischer M, Pinkert S, Mattheus W, Mainka A, Mürbe D. Formant frequencies and bandwidths of the vocal tract transfer function are affected by the mechanical impedance of the vocal tract wall. Biomech Model Mechanobiol 2015;14:719-33. doi: 10.1007/s10237-014-0632-2.
  34. Kılıç MA, Tüysüz O, Hanege FM, Paltura C. PraatAssisted Nasalance Meter: A Low-Cost Nasalance Measurement System for Evaluation of Nasal Resonance Disorders. Hamidiye Med J 2021;2:116-21.
  35. Boersma P, Weenink D, Shchupak A. Praat: doing phonetics by computer [computer program]. Version 7.0.02. Amsterdam: University of Amsterdam; 2026.
  36. Arai T. Vocal-tract models and their applications in education for intuitive understanding of speech production. Acoust Sci Technol 2016;37:148-56. doi: 10.1250/ ast.37.148.
  37. Klatt DH. Software for a cascade/parallel formant synthesizer. J Acoust Soc Am 1980;67:971-95.
  38. Titze IR, Story BH. Rules for controlling low-dimensional vocal fold models with muscle activation. J Acoust Soc Am 2002;112:1064-76. doi: 10.1121/1.1496080.
  39. Howe MS. Acoustics of fluid-structure interactions. Cambridge; New York: Cambridge University Press; 1998.
  40. Krane MH. Aeroacoustic production of low-frequency unvoiced speech sounds. J Acoust Soc Am 2005;118:410- 27. doi: 10.1121/1.1862251.
  41. Zheng X, Bielamowicz S, Luo H, Mittal R. A computational study of the effect of false vocal folds on glottal flow and vocal fold vibration during phonation. Ann Biomed Eng 2009;37:625-42. doi: 10.1007/s10439-008-9630-9.
  42. Teager HM, Teager SM. Evidence for nonlinear sound production mechanisms in the vocal tract. In: Hardcastle WJ, Marchal A, editors. Speech production and speech modelling. Dordrecht: Springer Netherlands; 1990. p. 241- 61. doi: 10.1007/978-94-009-2037-8_10.
  43. Chan RW, Titze IR. Viscoelastic shear properties of human vocal fold mucosa: Measurement methodology and empirical results. J Acoust Soc Am 1999;106:2008-21. doi: 10.1121/1.427947.
  44. Verdolini K, Titze IR, Fennell A. Dependence of phonatory effort on hydration level. J Speech Hear Res 1994;37:1001- 7. doi: 10.1044/jshr.3705.1001.