# Pose estimation and encode with TFLite
Source: [https://docs.qualcomm.com/doc/80-70014-50/topic/single-camera-stream-with-pose-estimation-and-encode.html](https://docs.qualcomm.com/doc/80-70014-50/topic/single-camera-stream-with-pose-estimation-and-encode.html)
The use cases use the PoseNet TFLite model to process a single camera stream with
pose estimation and encode the stream as a H.264 bitstream.
## **Variant 1: Use qtioverlay plugin to apply pose estimation overlay**
Use the following commands to execute the use
case:
setprop persist.overlay.use_c2d_blit 2Copy to clipboard
gst-launch-1.0 -e \
qtiqmmfsrc name=camsrc ! video/x-raw\(memory:GBM\),format=NV12,width=1280,height=720,framerate=30/1,compression=ubwc ! queue ! tee name=split \
split. ! queue ! qtimetamux name=metamux ! queue ! qtioverlay ! queue ! v4l2h264enc capture-io-mode=5 output-io-mode=5 ! h264parse ! queue ! mp4mux ! queue ! filesink location=/opt/video.mp4 \
split. ! queue ! qtimlvconverter ! queue ! qtimltflite delegate=external external-delegate-path=libQnnTFLiteDelegate.so external-delegate-options="QNNExternalDelegate,backend_type=htp;" model=/opt/posenet_mobilenet_v1.tflite ! queue ! qtimlvpose threshold=51.0 results=2 module=posenet labels=/opt/posenet_mobilenet_v1.labels constants="Posenet,q-offsets=<128.0,128.0,117.0>,q-scales=<0.0784313753247261,0.0784313753247261,1.3875764608383179>;" ! text/x-raw ! queue ! metamux.Copy to clipboard
To stop the use case, press CTRL + C.
Figure : Pipeline for pose estimation and encode using qtioverlay

The figure shows the flow of the use case execution:
1. Identify poses of people in scenes from video stream coming through camera
source.
2. Overlay the available poses using overlaylib.
3. Encode this stream as a H.264 bitstream.
4. Multiplex the stream in a MP4 container and store it as a MP4 file.
The table provides the sequential processing stages of the pipeline execution:
| Process | Description |
| --- | --- |
| [qtiqmmfsrc](https://docs.qualcomm.com/doc/80-70014-50/topic/qtiqmmfsrc.html) |
- Collects the video stream (source) and creates two copies of
the source:
- One stream is sent to qtimetamux plugin to retain
the video stream.
- The other stream is sent to a ML inferencing
pipeline.
|
| **Preprocessing** | **Preprocessing** |
| [qtimlvconverter](https://docs.qualcomm.com/doc/80-70014-50/topic/qtimlvconverter.html) |
- Receives the video stream on its sink pad.
- Performs preprocessing:
- Color conversion
- Scaling down/up
- Normalization on the stream data when model expects
floating point values as input
- Converts the video stream to a tensor stream on its source
pad.The PoseNet model uses this tensor stream for
inferencing.
|
| **Inferencing** | **Inferencing** |
| [qtimltflite](https://docs.qualcomm.com/doc/80-70014-50/topic/qtimltflite.html) |
- Loads the PoseNet model.
- Modifies the graph for the chosen delegate.
- Receives the tensor stream on its sinkpad.
- Executes the inference and produces tensor stream with the
pose estimation results on its source pad.
|
| **Postprocessing** | **Postprocessing** |
| [qtimlvpose](https://docs.qualcomm.com/doc/80-70014-50/topic/qtimlvpose.html) |
- Receives the inference tensors from a PoseNet model on its
sinkpad.
- Converts the tensors into formats such as video or text that
can be processed by the multimedia plugins later.
- Applies the threshold to the chosen number of results.
- Loads the corresponding modules of the pose estimation
models. In this use case, qtimlvpose does the
following:
- Loads the PoseNet submodule.
- Produces results as structures of text.
- Sends them to the sinkpad of qtimetamux.
|
| [qtimetamux](https://docs.qualcomm.com/doc/80-70014-50/topic/qtimetamux.html) |
- Receives the video stream and text stream with pose results
corresponding to video stream on its sinkpads.
- Produces GST buffers with contents of video stream on its
sink pad.
- Adds poses from data sinkpad to GST buffer meta (meta
muxing) on its source pad.
|
| [qtioverlay](https://docs.qualcomm.com/doc/80-70014-50/topic/qtioverlay.html) |
- Receives the multiplexed stream.
- Overlays the poses on the VideoFrame using CL.
- Produces GST buffers with overlays in its source pad.
|
| [v4l2h264enc](https://docs.qualcomm.com/doc/80-70014-50/topic/v4l2h264enc.html) |
- Applies parameters to each frame of the video stream it is
receiving on its sinkpad.
- Encodes it into bitstream and sends it over its
sourcepad.
|
| h264parse | Adds additional information corresponding to the bitstream to
GStreamer buffer meta. |
| mp4mux | Receives the buffers and creates containers format specification
buffers. |
| **Output** | **Output** |
| Filesink | Stores the resulting stream in a
/opt/video.mp4 file. |
| Playback | Use the following command to pull video.mp4
from the host machine and play it on a media player
application:
`scp root@ device>:/opt/ directory>` |
## Variant 2: Use qtivcomposer to mix original frame with pose estimation
mask
Use the following command to execute the use
case:
gst-launch-1.0 -e --gst-debug=2 \
qtiqmmfsrc name=camsrc ! video/x-raw\(memory:GBM\),format=NV12,width=1280,height=720,framerate=30/1,compression=ubwc ! queue ! tee name=split \
split. ! queue ! qtivcomposer name=mixer sink_1::dimensions="<1920,1080>" ! queue ! video/x-raw\(memory:GBM\),format=NV12,width=1920,height=1080,interlace-mode=progressive,colorimetry=bt601 ! v4l2h264enc capture-io-mode=5 output-io-mode=5 ! h264parse ! queue ! mp4mux ! queue ! filesink location=/opt/video.mp4 \
split. ! queue ! qtimlvconverter ! queue ! qtimltflite delegate=external external-delegate-path=libQnnTFLiteDelegate.so external-delegate-options="QNNExternalDelegate,backend_type=htp;" model=/opt/posenet_mobilenet_v1.tflite ! queue ! qtimlvpose threshold=51.0 results=2 module=posenet labels=/opt/posenet_mobilenet_v1.labels constants="Posenet,q-offsets=<128.0,128.0,117.0>,q-scales=<0.0784313753247261,0.0784313753247261,1.3875764608383179>;" ! video/x-raw,format=BGRA,width=640,height=360 ! queue ! mixer.Copy to clipboard
To stop the use case, press CTRL + C.
Figure : Pipeline for pose estimation and encode using qtivcomposer

The figure shows the flow of the use case execution:
- Classify scenes from the video stream coming through camera source.
- Compose the poses and video stream together using qtivcomposer.
- Encode this stream as a H.264 bitstream.
- Multiplex in a MP4 container and storing it as a MP4 file.
The table provides the sequential processing stages of the pipeline execution:
| Process | Description |
| --- | --- |
| [qtiqmmfsrc](https://docs.qualcomm.com/doc/80-70014-50/topic/qtiqmmfsrc.html) |
- Collects the video stream (source) and creates two copies of
the source:
- One stream is sent to qtivcomposer plugin to retain
the video stream.
- The other stream is sent to a ML inferencing
pipeline.
|
| **Preprocessing** | **Preprocessing** |
| [qtimlvconverter](https://docs.qualcomm.com/doc/80-70014-50/topic/qtimlvconverter.html) |
- Receives the video stream on its sink pad.
- Performs preprocessing:
- Color conversion
- Scaling down/up
- Normalization on the stream data when model expects
floating point values as input
- Converts the video stream to a tensor stream on its source
pad.The PoseNet model uses this tensor stream for
inferencing.
|
| **Inferencing** | **Inferencing** |
| [qtimltflite](https://docs.qualcomm.com/doc/80-70014-50/topic/qtimltflite.html) |
- Loads the PoseNet model.
- Modifies the graph for the chosen delegate.
- Receives the tensor stream on its sinkpad.
- Executes the inference and produces tensor stream with the
pose estimation results on its source pad.
|
| **Postprocessing** | **Postprocessing** |
| [qtimlvpose](https://docs.qualcomm.com/doc/80-70014-50/topic/qtimlvpose.html) |
- Receives the inference tensors from a PoseNet model on its
sinkpad.
- Converts the tensors into formats such as video or text that
can be processed by the multimedia plugins later.
- Applies the threshold to the chosen number of results.
- Loads the corresponding modules of the pose estimation
models. In this use case, qtimlvpose does the
following:
- Loads the PoseNet submodule.
- Produces results as video frames with poses
drawn.
- Sends them to sinkpad of the qtivcomposer.
|
| |
- Receives the original video stream and video stream of poses
on its sinkpads.
- On its sourcepad, produces the GST buffers with contents
composed of video streams from its sinkpads.
|
| [qtioverlay](https://docs.qualcomm.com/doc/80-70014-50/topic/qtioverlay.html) |
- Receives the multiplexed stream.
- Overlays the poses on the VideoFrame using CL.
- Produces GST buffers with overlays in its source pad.
|
| [v4l2h264enc](https://docs.qualcomm.com/doc/80-70014-50/topic/v4l2h264enc.html) |
- Applies parameters to each frame of the video stream it is
receiving on its sinkpad.
- Encodes it into bitstream and sends it over its
sourcepad.
|
| h264parse | Adds additional information corresponding to the bitstream to
GStreamer buffer meta. |
| mp4mux | Receives the buffers and creates containers format specification
buffers. |
| **Output** | **Output** |
| Filesink | Stores the resulting stream in a
/opt/video.mp4 file. |
| Playback | Use the following command to pull video.mp4
from the host machine and play it on a media player
application:
`scp root@ device>:/opt/ directory>` |
**Parent Topic:** [TensorFlow Lite use cases](https://docs.qualcomm.com/doc/80-70014-50/topic/tensorflow-lite-use-cases.html)
Last Published: Oct 27, 2025
[Previous Topic
Pose estimation and display with TFLite](https://docs.qualcomm.com/bundle/publicresource/80-70014-50/topics/single-camera-stream-with-pose-estimation-and-display.md) [Next Topic
Qualcomm Neural Processing SDK use cases](https://docs.qualcomm.com/bundle/publicresource/80-70014-50/topics/qualcomm-neural-processing-sdk-use-cases.md)