Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A mel filter bank converts each short-time speech spectrum into a smaller set of perceptually spaced frequency-band energies. The usual workflow is to frame and window the waveform, compute an STFT, apply overlapping triangular filters spaced on the mel scale, sum each band’s energy, and optionally take a logarithm. Log-mel features stop there; MFCCs apply a cepstral transform to the log-mel values.

What a mel filter bank does

A filter bank is a collection of frequency-selective filters that decomposes a signal into bands. In speech recognition, this provides a compact description of how energy is distributed across frequency, in a way intended to resemble the ear’s greater resolution at lower frequencies.

Mel filter banks normally use overlapping triangular windows. Their center frequencies are evenly spaced on a mel scale rather than on the ordinary hertz scale. For each short-time spectrum frame, the filters weight neighboring FFT bins, and the weighted values are summed into one number per filter. The result is a time-by-filter matrix: one mel value for every filter in every frame.

Converting a spectrogram to mel features

  1. Frame the waveform. Split the audio into short, usually overlapping windows so that speech is approximately stationary within each frame.
  2. Apply a window function. A Hamming window is a common choice before the frequency transform.
  3. Compute a spectrum. Use an STFT or another frequency-domain representation. Decide whether the filters will receive magnitude or power values.
  4. Construct the mel filters. Map the permitted frequency range to mel points and create overlapping triangular filters. Each triangle rises to a peak, typically 1.0, then falls to zero at the neighboring points.
  5. Aggregate each band. Multiply the spectrum by every filter and sum the weighted bins. This gives one energy value per mel band and frame.
  6. Compress the dynamic range when needed. Apply a natural or base-10 logarithm, or convert to decibels, to obtain log-mel features.

A 2020 methods paper used 40 ms windows extracted every 10 ms, a Hamming-windowed STFT, 128 triangular mel filters, and a logarithm. Those are that study’s experimental settings, not universal defaults.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Shure MVX2U Gen 2 XLR-to-USB-C Audio Interface
  • HIGH-PERFORMANCE XLR-TO-USB-C INTERFACE - Streamline your recording and streaming setups on desktop, tablet, or smartphone with clean, consistent audio across devices using any connected XLR microphone.
  • ADVANCED AUDIO PROCESSING - Features onboard Shure Digital Audio Processing including Auto Level Mode, Real-Time Denoiser, and Digital Popper Stopper for zero-latency audio with any XLR microphone.
  • AUTO LEVEL MODE - Automatically adjusts gain in real time with onboard DSP for consistent output. Choose your preferred tone from Dark, Natural, or Bright for tailored audio performance.
  • PLUG-AND-PLAY CONVENIENCE - Instantly convert any dynamic or condenser XLR mic for professional podcasting or livestreaming. Provides up to +60 dB clean gain and 48V phantom power for your microphone.
  • MOTIV APP COMPATIBILITY - Manage settings on desktop, smartphone, or tablet using MOTIV Mix, MOTIV Audio, and MOTIV Video apps. Activate audio processing, customize sound with tone, EQ, compression, and limiter for professional results.

Why frequencies are measured in mels

The mel scale is perceptual: it allocates relatively more resolution to lower frequencies and compresses spacing as frequency rises. Two frequencies that are far apart in hertz at the high end may occupy a smaller distance on the mel scale than the same hertz difference at the low end. This can reduce redundant high-frequency detail while preserving distinctions that are more important for many speech tasks.

Mel formulas are not interchangeable

There is no single implementation-independent mel conversion. NVIDIA documentation exposes both a Slaney option and an HTK option. The HTK formula is:

m = 2595 × log10(1 + f / 700)

where f is frequency in hertz and m is the mel value. The Slaney approach is linear below 1 kHz and logarithmic above it. Selecting a different formula changes filter center frequencies and therefore the feature tensor. Record the formula whenever you need reproducible results.

Mel spectrogram, log-mel, and MFCCs

Representation How it is produced What it contains
Mel spectrogram Apply a mel filter bank to each spectrum frame. Band-aggregated magnitude or power values on a perceptual frequency axis.
Log-mel feature Take the logarithm (or a dB conversion) of mel-band values. Compressed band energies, often easier for a model to handle across a large dynamic range.
MFCC Apply a cepstral transform after creating the log-mel representation. Cepstral coefficients that summarize the spectral envelope rather than retaining the mel bands directly.

NVIDIA’s audio example treats MFCCs as an alternative representation derived from a mel-frequency spectrogram and shows the sequence of spectrogram, mel filter bank, decibel conversion, and MFCC computation. MFCCs are therefore not a different way to place the triangular filters; they are a further transformation of log-mel information.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parameters that determine the feature tensor

Two systems can both report “mel spectrogram” and still produce incompatible arrays. The important choices are:

Parameter Effect What to record
Number of filters Controls frequency resolution and feature width. For example, 24, 40, 80, or 128 bands.
Lower and upper frequency Defines which part of the spectrum is represented. Exact minimum and maximum in hertz.
Sample rate Sets the Nyquist limit and the mapping from FFT bins to hertz. Samples per second and any resampling step.
FFT size Sets the spacing of the discrete frequency bins. FFT length and whether zero-padding is used.
Window and hop Sets time resolution, overlap, and number of frames. Window length, hop length, window type, and units.
Filter shape and overlap Determines how neighboring bins contribute to each band. Triangular design, spacing rule, and peak normalization.
Mel formula Changes filter center and edge frequencies. HTK, Slaney, or the toolkit’s documented alternative.
Input quantity Changes scale before aggregation. Magnitude, power, log power, or dB.
Normalization Changes each filter’s amplitude weighting. Area, peak, or toolkit-specific normalization.

TensorFlow’s linear_to_mel_weight_matrix maps linear frequencies from 0 through half the sample rate into a chosen number of mel bins with triangular weights whose peaks are 1.0. MathWorks documents half-overlapped triangular filters equally spaced on the mel scale and exposes frequency range, number of bands, and normalization controls. NVIDIA DALI exposes filter count, frequency limits, sample rate, formula, and normalization.

How many mel filters should you use?

There is no universal optimum. More filters preserve finer spectral detail but increase feature width and computation; fewer filters compress the spectrum more aggressively and may discard distinctions useful to a particular task.

Rank #2
PUPGSIS Gaming Audio Mixer for PC Streaming, Soundboard with Voice Changer
  • This sound card is not compatible with 48V dynamic microphones or USB microphones. It only supports XLR microphones. (Note: Connecting an XLR microphone requires a 1/4" TRS to XLR cable, which is available as part of a promotional offer and must be added separately.)
  • All-in-One Audio Interface for Streaming – This mixer works as a complete audio hub for live streaming, podcasting, and gaming. It features a 1/4" TRS dynamic microphone input, built-in reverb, 4 custom sound effects pads, and a voice changer, so you can enhance your voice and engage your audience with creative audio in real time.
  • Effective Noise Cancellation – Equipped with advanced noise reduction technology, the PUPGSIS mixer filters out background hum, fan noise, and other unwanted sounds. Your viewers will hear only your clear, professional voice – ideal for noisy gaming rooms or home studios.
  • Customizable Sound Effects & Voice Changer – Personalize your stream with 4 programmable sound effect buttons. Load your own audio clips (laugh tracks, claps, alarms, etc.) and activate them instantly. The built‑in voice changer lets you alter your pitch for fun character voices or anonymous commentary.
  • Adjustable Reverb for Professional Vocals – The mixer features a fully adjustable reverb effect, allowing you to dial in exactly the right amount of room ambience for your voice. Whether you want a subtle studio echo or a dramatic live‑stage sound, the dedicated reverb control lets you fine‑tune it on the fly – no software needed.
  • 24 filters: a compact configuration; ISIP shows an example using 24 triangular filters at an 8 kHz sample frequency.
  • 40 filters: a common-sized feature representation when a moderate input width is desired.
  • 80 filters: a higher-resolution option often suitable when the model and data support wider inputs.
  • 128 filters: used in the cited 2020 study; NVIDIA DALI’s documented nfilter default is 128 in archived version 1.41.0, with a documented default sample rate of 44,100 Hz for that operator page.

Treat these as design points, not guarantees of accuracy. Choose a count that matches the sample rate, frequency range, model input shape, and amount of training data, then validate it on the target task. Keep every other preprocessing choice fixed when comparing counts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Practical conversion checklist

  1. Set or verify the audio sample rate and decide whether to resample.
  2. Choose window and hop durations appropriate to the speech timing you need.
  3. Select an FFT length and determine whether the model consumes magnitude or power.
  4. Set the minimum and maximum frequencies; do not silently assume the full Nyquist range.
  5. Choose the filter count and mel formula.
  6. Generate overlapping triangular filters and document their normalization.
  7. Multiply each frame’s spectrum by the filter bank and sum along frequency.
  8. Apply log or dB compression consistently, including the policy for zero or near-zero values.
  9. Store the resulting tensor shape and axis order with the model configuration.

Troubleshooting mismatched mel features

Same audio, different tensor shape

Check frame length, hop length, sample rate, FFT size, and filter count first. A different hop changes the number of time frames; a different filter count changes the feature dimension.

Same shape, different numerical values

Compare the mel formula, frequency limits, normalization, triangular filter construction, and whether the input was magnitude or power. Also check whether logarithm or decibel conversion occurred before or after filtering.

Unexpected high- or low-frequency behavior

Verify the Nyquist limit implied by the sample rate and the explicitly configured upper frequency. A toolkit default can be version-specific; NVIDIA DALI’s documented defaults cited above are tied to archived version 1.41.0.

MFCCs do not match log-mel features

That is expected: MFCC extraction adds a cepstral transform after log-mel processing. Compare the same intermediate log-mel matrix before investigating later coefficients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Key takeaway

Mel filtering is the frequency-band aggregation stage applied independently to each short-time spectrum frame. The triangular filters, mel formula, frequency range, time-frequency resolution, normalization, and logarithmic compression are all part of the feature definition. For reproducible machine-learning experiments, save those settings—not just the label “mel spectrogram.”

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.