# Audio classification decode and display with LiteRT
Source: [https://docs.qualcomm.com/doc/80-70018-50/topic/audio-classification-with-litert.html](https://docs.qualcomm.com/doc/80-70018-50/topic/audio-classification-with-litert.html)
The use cases use the YAMNet LiteRT model to classify and decode audio samples from a
microphone and a file source.
To stop the use cases, use CTRL + C.
## Audio classification on audio samples from microphone
Run this use
case:
gst-launch-1.0 -v pulsesrc ! audio/x-raw,format=S16LE ! audiobuffersplit output-buffer-size=31200 ! \
qtimlaconverter sample-rate=16000 feature=lmfe params="params,nfft=512,nhop=160,nmels=64,chunklen=0.96;" ! queue ! \
qtimltflite model=/etc/models/yamnet.tflite ! qtimlaclassification module=yamnet labels=/etc/labels/yamnet.labels ! \
video/x-raw,width=640,height=360 ! queue ! waylandsink sync=false fullscreen=trueCopy to clipboard
The figure shows the flow of the use case execution:
Figure : Pipeline flow for audio classification and display

The table provides the sequential processing stages of the pipeline execution:
Table : Pipeline processing stages for audio classification
| Process | Description |
| --- | --- |
| [pulsesrc](https://docs.qualcomm.com/doc/80-70018-50/topic/pulsesrc.html) | Collects the audio stream (source) from the microphone. |
| audiobuffersplit | Splits the incoming audio buffers into equal sized
chunks. |
| [qtimlaconverter](https://docs.qualcomm.com/doc/80-70018-50/topic/qtimlaconverter.html) | Performs preprocessing on the audio stream and converts the
stream to a tensor stream.
The audio classification model uses
this tensor stream for inferencing. |
| [qtimltflite](https://docs.qualcomm.com/doc/80-70018-50/topic/qtimltflite.html) |
- Loads the model.
- Modifies the graph for the chosen delegate.
- Receives the tensor stream on its sinkpad.
- Runs the inference and produces a tensor stream with the
inference results on its source pad.
|
| [qtimlaclassification](https://docs.qualcomm.com/doc/80-70018-50/topic/qtimlaclassification.html) | Handles the audio classification inference results:
- Applies a threshold to the chosen number of results.
- Creates text overlay for classes.
|
| [Waylandsink](https://docs.qualcomm.com/doc/80-70018-50/topic/waylandsink.html) |
- Receives the video and audio streams on its sinkpad.
- Submits the streams to Weston.
- Weston renders the video stream and the classified audio
generated for that scene on a local display device.
|
## Audio classification and decode on audio samples from file source
- Run the use case using a FLAC
decoder:
gst-launch-1.0 -e --gst-debug=2 filesrc location=/etc/media/video_FLAC.mp4 ! qtdemux name=demux demux. ! queue ! h264parse ! \
v4l2h264dec capture-io-mode=4 output-io-mode=4 ! video/x-raw, format=NV12 ! qtivcomposer name=mixer sink_1::position="<50, 50>" sink_1::dimensions="<368, 64>" ! \
queue ! waylandsink fullscreen=true demux. ! queue ! flacparse ! flacdec ! queue ! audioconvert ! audio/x-raw,format=S16LE ! audioresample ! \
audiobuffersplit output-buffer-size=31200 ! queue ! qtimlaconverter sample-rate=16000 feature=lmfe params="params,nfft=512,nhop=160,nmels=64,chunklen=0.96;" ! \
queue ! qtimltflite name=infeng model=/etc/models/yamnet.tflite ! qtimlaclassification name=postproc threshold=10.0 results=3 module=yamnet \
labels=/etc/labels/yamnet.labels ! video/x-raw,format=BGRA,width=368,height=64 ! queue ! mixer.Copy to clipboard
Figure : Pipeline flow for audio classification and decode–FLAC
decoder

- Run the use case using the mpg123audiodec
decoder:
gst-launch-1.0 -e --gst-debug=2 filesrc location=/etc/media/video_mp3.mp4 ! qtdemux name=demux demux. ! queue ! h264parse ! v4l2h264dec capture-io-mode=4 output-io-mode=4 ! video/x-raw, format=NV12 ! qtivcomposer name=mixer sink_1::position="<50, 50>" sink_1::dimensions="<368, 64>" ! queue ! waylandsink fullscreen=true demux. ! queue ! mpegaudioparse ! mpg123audiodec ! queue ! audioconvert ! audio/x-raw,format=S16LE ! audioresample ! audiobuffersplit output-buffer-size=31200 ! queue ! qtimlaconverter sample-rate=16000 feature=lmfe params="params,nfft=512,nhop=160,nmels=64,chunklen=0.96;" ! queue ! qtimltflite name=infeng model=/etc/models/yamnet.tflite ! qtimlaclassification name=postproc threshold=10.0 results=3 module=yamnet labels=/etc/labels/yamnet.labels ! video/x-raw,format=BGRA,width=368,height=64 ! queue ! mixer.Copy to clipboard
Figure : Pipeline flow for audio classification and decode–mpg123audiodec
decoder

The table provides the sequential processing stages of the pipeline execution:
Table : Pipeline processing stages for audio classification using FLAC and
mpg123audiodec decoders
| Plugin | Description |
| --- | --- |
| File source: filesrc | Captures the video and audio stream, followed by qtdemux, which
demultiplexes the stream. |
| Audio and video capsfilter | Ensures that the video and audio streams are in the correct
format. |
|
- Audio: mpegaudioparse or flacparse
- Video: h264parse
| Parses the audio and the H.264 video. |
| | Decodes the audio and video. |
| audioconvert | Converts the audio buffers between various possible
formats. |
| audioresample | Resamples the audio buffers to different sample rates. |
| audiobuffersplit | Splits the incoming audio buffers into equal sized
chunks. |
| [qtimlaconverter](https://docs.qualcomm.com/doc/80-70018-50/topic/qtimlaconverter.html) | Performs preprocessing on the audio stream and converts the
stream to a tensor stream.
The audio classification model uses
this tensor stream for inferencing. |
| [qtimltflite](https://docs.qualcomm.com/doc/80-70018-50/topic/qtimltflite.html) | Performs inferencing using the YAMNet model. |
| [qtimlaclassification](https://docs.qualcomm.com/doc/80-70018-50/topic/qtimlaclassification.html) | Handles the audio classification inference results:
- Applies a threshold to the chosen number of results.
- Creates text overlay for classes.
|
| [qtivcomposer](https://docs.qualcomm.com/doc/80-70018-50/topic/qtivcomposer.html) | Combines the text overlay for classification results and video
preview. |
| [Waylandsink](https://docs.qualcomm.com/doc/80-70018-50/topic/waylandsink.html) |
- Waylandsink submits the video stream received on its sink
pad to Weston.
- Weston renders the video stream on a local display.
|
**Parent Topic:** [LiteRT use cases](https://docs.qualcomm.com/doc/80-70018-50/topic/tensorflow-lite-use-cases.html)
**Related Resources**
- [Audio classification](https://docs.qualcomm.com/doc/80-70018-50/topic/audio-classification.html)
Last Published: Jan 30, 2026
[Previous Topic
Image classification and encode with LiteRT](https://docs.qualcomm.com/bundle/publicresource/80-70018-50/topics/single-camera-stream-with-image-classification-and-encode.md) [Next Topic
Object detection and display with LiteRT](https://docs.qualcomm.com/bundle/publicresource/80-70018-50/topics/single-camera-stream-with-object-detection-and-display.md)