# LLM Pipeline
LLM-specific pipeline implementation, configuration, and context classes.
## LLMPipeline
- *class* qairt.experimental.pipeline.torch.llm.pipeline.LLMPipeline(*config: [LLMPipelineConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.LLMPipelineConfig)*)
- Bases: [`Pipeline`](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-base.html#qairt.experimental.pipeline.torch.common.bases.pipeline.Pipeline)[[`LLMPipelineConfig`](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.LLMPipelineConfig)]
Pipeline for PyTorch LLMs.
- *property* config*: [LLMPipelineConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.LLMPipelineConfig)*
- Return the domain-specific pipeline configuration.
- evaluate(*\*\*kwargs*) → dict[str, float]
- Run the configured evaluation metrics on the last stage’s output.
Delegates to the last stage’s `evaluate()`. A caller may pass
`evaluator_config` via `**kwargs` to override the pipeline-level config.
- Returns
- `{display_name: score}` e.g. `{"PPL_wikitext": 8.41}`.
- Raises
- - **RuntimeError** – If `construct()` has not been called, or the last
stage produces no evaluable model.
- **ValueError** – If no `evaluator_config` or metrics are configured.
- *classmethod* from\_pretrained(*model\_id\_or\_path: str*, *recipe: Optional[Union[str, Path, dict[str, Any]]] = None*, *\*\*kwargs: Any*) → [LLMPipeline](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.LLMPipeline)
- Create a pipeline from a recipe and execute the `model_loader` stage.
- Parameters
- - **model\_id\_or\_path** – HuggingFace model ID or local path.
- **recipe** – Recipe file path or dict.
- **\*\*kwargs** – Forwarded to the base `from_pretrained` as config overrides.
- Returns
- An `LLMPipeline` instance with the first stage executed.
- generate(*prompt: Union[str, list[dict[str, str]], Path]*, *device: Optional[Any] = None*, *\*\*kwargs*) → TextGenerationResult
- Generate text using the last stage in the pipeline.
The last stage must have `can_generate = True` and implement `generate()`.
Users should be aware of the generation capabilities of the last stage and
provide appropriate kwargs (e.g. `max_length`, `num_beams`).
- Parameters
- - **prompt** – Input prompt for generation. Can be a plain string, a list of
chat messages (dicts with “role” and “content” keys), or a Path to
a JSON file containing chat messages.
- **device** – Optional device specification (forwarded to the stage).
- **\*\*kwargs** – Additional generation parameters forwarded to the last stage.
- Returns
- Generation result from the last stage. Stages before `genai_builder`
(e.g. `model_loader`, `quantization`) return plain decoded text,
which is wrapped in a `TextGenerationResult` for a consistent return type.
- Raises
- - **RuntimeError** – If `construct()` has not been called or the last stage
has not been executed.
- **NotImplementedError** – If the last stage does not support generation.
## LLMPipelineConfig
- *class* qairt.experimental.pipeline.torch.llm.pipeline.LLMPipelineConfig(*\**, *model\_id\_or\_path: str*, *cache\_dir: str = './workspace'*, *backend: [BackendType](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-api-configs.html#qairt.api.configs.common.BackendType) = BackendType.HTP*, *soc\_details: Optional[Union[qti.aisw.tools.core.utilities.devices.api.device\_definitions.SocDetails, str]] = None*, *log\_level: Optional[str] = None*, *task: str = 'text-generation'*, *features: [PipelineFeatures](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.PipelineFeatures) = None*, *enable\_cache: bool = False*, *checkpoint: Optional[str] = None*, *enable\_observers: bool = False*, *observers: dict[str, Any] = None*, *stage\_info: OrderedDict[str, dict[str, Any]] = None*, *custom\_stage: list[typing.Annotated[pathlib.Path, PathType(path\_type='file')]] = None*, *generator\_config: [qairt.experimental.pipeline.torch.common.configs.GeneratorConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.GeneratorConfig) | None = None*, *evaluator\_config: [qairt.experimental.pipeline.torch.common.configs.EvaluatorConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.EvaluatorConfig) | None = None*, *exporter\_config: [qairt.experimental.pipeline.torch.common.configs.ExporterConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.ExporterConfig) | None = None*, *eaglet: dict[str, Any] | None = None*)
- Bases: [`PipelineConfig`](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-base.html#qairt.experimental.pipeline.torch.common.bases.pipeline.PipelineConfig), [`LLMPipelineContext`](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.LLMPipelineContext)
Global pipeline configuration for LLM pipelines.
- add\_config(*stage\_name: str*, *config: [qairt.experimental.pipeline.torch.common.configs.ExporterConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.ExporterConfig) | [qairt.experimental.pipeline.torch.common.configs.GeneratorConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.GeneratorConfig) | [qairt.experimental.pipeline.torch.common.configs.EvaluatorConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.EvaluatorConfig)*) → None
- Adds a config for a specific stage.
- Parameters
- - **stage\_name** – Name of the stage to associate the config with.
- **config** – Must be one of ExporterConfig, GeneratorConfig or EvaluatorConfig.
Stage-specific configs override global configs for the same stage.
- add\_dataloader(*stage\_name: str*, *dataloader: DataLoader*) → None
- Adds a data loader for a specific stage.
- Parameters
- - **stage\_name** – Name of the stage to associate the dataloader with (e.g.,
“quantization” for calibration dataloader)
- **dataloader** – The DataLoader instance to add
- *field* eaglet*: dict[str, Any] | None* *= None*
- - Validated by
- - `_check_unsupported_features`
- *field* evaluator\_config*: [EvaluatorConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.EvaluatorConfig) | None* *= None*
- - Validated by
- - `_check_unsupported_features`
- *field* exporter\_config*: [ExporterConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.ExporterConfig) | None* *= None*
- - Validated by
- - `_check_unsupported_features`
- *classmethod* from\_recipe(*recipe: Union[str, Path, dict[str, Any]]*) → [LLMPipelineConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.LLMPipelineConfig)
- Load pipeline configuration from a YAML recipe file or dict.
- Parameters
- **recipe** – Path to YAML recipe file, or a pre-loaded recipe dict.
- Returns
- `LLMPipelineConfig` instance.
- *field* generator\_config*: [GeneratorConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.GeneratorConfig) | None* *= None*
- - Validated by
- - `_check_unsupported_features`
- model\_computed\_fields*: ClassVar[dict[str, ComputedFieldInfo]]* *= {}*
- A dictionary of computed field names and their corresponding ComputedFieldInfo objects.
- model\_config*: ClassVar[ConfigDict]* *= {'arbitrary\_types\_allowed': True, 'protected\_namespaces': ()}*
- Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- model\_fields*: ClassVar[dict[str, FieldInfo]]* *= {'backend': FieldInfo(annotation=BackendType, required=False, default=<BackendType.HTP: 'HTP'>), 'cache\_dir': FieldInfo(annotation=str, required=False, default='./workspace'), 'checkpoint': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'custom\_stage': FieldInfo(annotation=list[Annotated[Path, PathType]], required=False, default\_factory=list), 'eaglet': FieldInfo(annotation=Union[dict[str, Any], NoneType], required=False, default=None), 'enable\_cache': FieldInfo(annotation=bool, required=False, default=False), 'enable\_observers': FieldInfo(annotation=bool, required=False, default=False), 'evaluator\_config': FieldInfo(annotation=Union[EvaluatorConfig, NoneType], required=False, default=None), 'exporter\_config': FieldInfo(annotation=Union[ExporterConfig, NoneType], required=False, default=None), 'features': FieldInfo(annotation=PipelineFeatures, required=False, default\_factory=PipelineFeatures), 'generator\_config': FieldInfo(annotation=Union[GeneratorConfig, NoneType], required=False, default=None), 'log\_level': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'model\_id\_or\_path': FieldInfo(annotation=str, required=True), 'observers': FieldInfo(annotation=dict[str, Any], required=False, default\_factory=dict), 'soc\_details': FieldInfo(annotation=Union[SocDetails, str, NoneType], required=False, default=None), 'stage\_info': FieldInfo(annotation=OrderedDict[str, dict[str, Any]], required=False, default\_factory=OrderedDict, alias\_priority=2, serialization\_alias='stages'), 'task': FieldInfo(annotation=str, required=False, default='text-generation')}*
- Metadata about the fields defined on the model,
mapping of field names to [FieldInfo][pydantic.fields.FieldInfo].
This replaces Model.__fields__ from Pydantic V1.
- model\_post\_init(*context: Any*, */*) → None
- We need to both initialize private attributes and call the user-defined model\_post\_init
method.
## LLMPipelineContext
- *class* qairt.experimental.pipeline.torch.llm.pipeline.LLMPipelineContext(*\**, *model\_id\_or\_path: str*, *cache\_dir: str = './workspace'*, *backend: [BackendType](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-api-configs.html#qairt.api.configs.common.BackendType) = BackendType.HTP*, *soc\_details: Optional[Union[qti.aisw.tools.core.utilities.devices.api.device\_definitions.SocDetails, str]] = None*, *log\_level: Optional[str] = None*, *task: str = 'text-generation'*, *features: [PipelineFeatures](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.PipelineFeatures) = None*)
- Bases: [`PipelineContext`](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-base.html#qairt.experimental.pipeline.torch.common.bases.pipeline_context.PipelineContext)
LLM-specific stage-visible pipeline context.
Extends `PipelineContext` with fields specific to large language models.
Injected into `StageConfig._pipeline_context` for all LLM stages.
- *field* features*: [PipelineFeatures](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.PipelineFeatures)* *[Optional]*
-
- model\_computed\_fields*: ClassVar[dict[str, ComputedFieldInfo]]* *= {}*
- A dictionary of computed field names and their corresponding ComputedFieldInfo objects.
- model\_config*: ClassVar[ConfigDict]* *= {'arbitrary\_types\_allowed': True, 'protected\_namespaces': ()}*
- Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- model\_fields*: ClassVar[dict[str, FieldInfo]]* *= {'backend': FieldInfo(annotation=BackendType, required=False, default=<BackendType.HTP: 'HTP'>), 'cache\_dir': FieldInfo(annotation=str, required=False, default='./workspace'), 'features': FieldInfo(annotation=PipelineFeatures, required=False, default\_factory=PipelineFeatures), 'log\_level': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'model\_id\_or\_path': FieldInfo(annotation=str, required=True), 'soc\_details': FieldInfo(annotation=Union[SocDetails, str, NoneType], required=False, default=None), 'task': FieldInfo(annotation=str, required=False, default='text-generation')}*
- Metadata about the fields defined on the model,
mapping of field names to [FieldInfo][pydantic.fields.FieldInfo].
This replaces Model.__fields__ from Pydantic V1.
- model\_post\_init(*context: Any*, */*) → None
- We need to both initialize private attributes and call the user-defined model\_post\_init
method.
- *field* task*: str* *= 'text-generation'*
-
## PipelineFeatures
- *class* qairt.experimental.pipeline.torch.llm.pipeline.PipelineFeatures(*\**, *speculative\_decoding: Tuple[bool, str | None] = (False, None)*, *lora: Optional[[LoRAFeatureConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-lora.html#qairt.experimental.pipeline.torch.llm.lora.configs.LoRAFeatureConfig)] = None*)
- Bases: `BaseModel`
Cross-cutting features that can be enabled/disabled for the entire pipeline.
These features can affect multiple stages. If set:
- stages will check to ensure feature related configs are provided.
- stages will enable feature-specific logic during execution.
**Field declaration order is load-bearing**: GeneratorFactory composes
generator mixins in the order fields are declared here. The first-declared
field gets the highest MRO position and therefore runs first in forward().
Do not reorder fields without updating the MRO contract test.
- feature\_keys() → list[str]
- Derive all active feature keys in field declaration order.
Iterates over the model fields in the order they are declared and
collects a key for every enabled field. For tuple-valued fields (e.g.
`speculative_decoding = (True, "eaglet")`) the key is
`"{field_name}_{method}"`; plain `bool` or `Optional` fields use
the field name directly.
**The returned order is load-bearing**: it determines mixin MRO
position inside `GeneratorFactory._compose()`. The first key in
the list produces the highest mixin in the MRO (called first in
`forward()`). Reordering fields in `PipelineFeatures` therefore
changes execution order — do not do so without updating the MRO
contract test.
- Returns
- Ordered list of active feature key strings, in field declaration
order. Empty when no feature is enabled.
- *field* lora*: Optional[[LoRAFeatureConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-lora.html#qairt.experimental.pipeline.torch.llm.lora.configs.LoRAFeatureConfig)]* *= None*
-
- model\_computed\_fields*: ClassVar[dict[str, ComputedFieldInfo]]* *= {}*
- A dictionary of computed field names and their corresponding ComputedFieldInfo objects.
- model\_config*: ClassVar[ConfigDict]* *= {}*
- Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- model\_fields*: ClassVar[dict[str, FieldInfo]]* *= {'lora': FieldInfo(annotation=Union[LoRAFeatureConfig, NoneType], required=False, default=None), 'speculative\_decoding': FieldInfo(annotation=Tuple[bool, Union[str, NoneType]], required=False, default=(False, None))}*
- Metadata about the fields defined on the model,
mapping of field names to [FieldInfo][pydantic.fields.FieldInfo].
This replaces Model.__fields__ from Pydantic V1.
- *field* speculative\_decoding*: Tuple[bool, str | None]* *= (False, None)*
-
## EagletV2Config
- *class* qairt.experimental.pipeline.torch.llm.eaglet.config.EagletV2Config(*\**, *eaglet\_version: str = 'v2'*, *draft\_safetensors\_path: Optional[str] = None*, *num\_draft\_layers: int = 1*, *model\_config\_overrides: Dict[str, Any] = None*, *num\_calibration\_batches: int = 128*, *calibration\_dataset: str = 'wikitext'*, *sequence\_length: int = 2048*, *dual\_alignment: bool = True*, *encoding\_tie\_recipe\_overrides: Optional[Dict[str, str]] = None*, *dm\_quantsim\_config\_file: Optional[str] = None*, *dsp\_arch: Optional[str] = None*, *dm\_filename\_prefix: str = 'eaglet\_dm'*, *export\_inputs\_embeds: bool = False*, *capture\_hidden\_states: bool = True*, *hidden\_states\_cache\_dir: Optional[str] = None*, *trimmed\_vocabulary\_length: Optional[int] = None*, *trimmed\_vocabulary\_path: Optional[str] = None*, *draft\_top\_k: int = 8*, *draft\_depth: int = 5*, *total\_draft\_tokens: int = 31*, *draft\_preselect: Optional[int] = None*, *draft\_temperature: Optional[float] = None*, *\*\*extra\_data: Any*)
- Bases: `BaseModel`
Configuration for Eaglet v2 speculative decoding pipeline.
Controls draft model construction, calibration, and encoding-tie behavior.
Used by the dm\_model\_loading and dm\_quantization stages.
The draft spec family is resolved automatically from the TM’s
`config.model_type` via `MODEL_TYPE_TO_FAMILY`. Users do not
need to specify it.
- eaglet\_version
- Eaglet version string (e.g. “v2”, “v2\_Qwen35”).
- Type
- str
- draft\_safetensors\_path
- Path to pretrained DM weights. None = random init.
- Type
- Optional[str]
- num\_draft\_layers
- Number of eaglet decoder layers in the draft model.
- Type
- int
- num\_calibration\_batches
- Number of batches for DM calibration.
- Type
- int
- calibration\_dataset
- Dataset name for calibration data.
- Type
- str
- sequence\_length
- Calibration dataloader context length (sequence length).
- Type
- int
- dual\_alignment
- Whether to run cross\_model\_tie for encoding alignment.
- Type
- bool
- encoding\_tie\_recipe\_overrides
- Override default ONNX op names for
encoding tie. Keys must match EncodingTieRecipe field names.
- Type
- Optional[Dict[str, str]]
- dm\_quantsim\_config\_file
- Path to HTP quantsim config JSON for
per-channel weight quantization. None = use default AIMET config.
- Type
- Optional[str]
- dsp\_arch
- DSP architecture version for HTP config selection (e.g. “v79”,
“v81”). None = use pipeline recipe value or CalibratorParams default.
- Type
- Optional[str]
- dm\_filename\_prefix
- Prefix for exported DM artifact files.
- Type
- str
- export\_inputs\_embeds
- Export the DM with `inputs_embeds` instead of
`input_ids` plus an in-graph embedding lookup.
- Type
- bool
- capture\_hidden\_states
- Whether TM quantization should capture hidden
states for DM calibration.
- Type
- bool
- hidden\_states\_cache\_dir
- Directory to store captured TM hidden states.
None = use pipeline cache\_dir.
- Type
- Optional[str]
- draft\_top\_k
- Number of top-k candidates during tree drafting.
- Type
- int
- draft\_depth
- Maximum tree depth for speculative drafting.
- Type
- int
- total\_draft\_tokens
- Maximum number of draft tokens per step.
- Type
- int
- draft\_preselect
- Preselect count for the two-stage token selection
(top-k → softmax → top-k). None uses log-softmax → top-k.
- Type
- Optional[int]
- draft\_temperature
- Softmax temperature for the two-stage selection.
Ignored when `draft_preselect` is None.
- Type
- Optional[float]
- model\_config\_overrides
- DM HF config overrides keyed by attribute name
(e.g. `{"hidden_size": 1024}`). Takes priority over the pretrained
DM config.json and the TM value. Unknown keys are set verbatim.
- Type
- Dict[str, Any]
- *field* calibration\_dataset*: str* *= 'wikitext'*
-
- *field* capture\_hidden\_states*: bool* *= True*
-
- *field* dm\_filename\_prefix*: str* *= 'eaglet\_dm'*
-
- *field* dm\_quantsim\_config\_file*: Optional[str]* *= None*
-
- *field* draft\_depth*: int* *= 5*
-
- *field* draft\_preselect*: Optional[int]* *= None*
-
- *field* draft\_safetensors\_path*: Optional[str]* *= None*
-
- *field* draft\_temperature*: Optional[float]* *= None*
-
- *field* draft\_top\_k*: int* *= 8*
-
- *field* dsp\_arch*: Optional[str]* *= None*
-
- *field* dual\_alignment*: bool* *= True*
-
- *field* eaglet\_version*: str* *= 'v2'*
-
- *field* encoding\_tie\_recipe\_overrides*: Optional[Dict[str, str]]* *= None*
-
- *field* export\_inputs\_embeds*: bool* *= False*
-
- *field* hidden\_states\_cache\_dir*: Optional[str]* *= None*
-
- model\_computed\_fields*: ClassVar[dict[str, ComputedFieldInfo]]* *= {}*
- A dictionary of computed field names and their corresponding ComputedFieldInfo objects.
- model\_config*: ClassVar[ConfigDict]* *= {'extra': 'allow', 'protected\_namespaces': ()}*
- Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- *field* model\_config\_overrides*: Dict[str, Any]* *[Optional]*
-
- model\_fields*: ClassVar[dict[str, FieldInfo]]* *= {'calibration\_dataset': FieldInfo(annotation=str, required=False, default='wikitext'), 'capture\_hidden\_states': FieldInfo(annotation=bool, required=False, default=True), 'dm\_filename\_prefix': FieldInfo(annotation=str, required=False, default='eaglet\_dm'), 'dm\_quantsim\_config\_file': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'draft\_depth': FieldInfo(annotation=int, required=False, default=5), 'draft\_preselect': FieldInfo(annotation=Union[int, NoneType], required=False, default=None), 'draft\_safetensors\_path': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'draft\_temperature': FieldInfo(annotation=Union[float, NoneType], required=False, default=None), 'draft\_top\_k': FieldInfo(annotation=int, required=False, default=8), 'dsp\_arch': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'dual\_alignment': FieldInfo(annotation=bool, required=False, default=True), 'eaglet\_version': FieldInfo(annotation=str, required=False, default='v2'), 'encoding\_tie\_recipe\_overrides': FieldInfo(annotation=Union[Dict[str, str], NoneType], required=False, default=None), 'export\_inputs\_embeds': FieldInfo(annotation=bool, required=False, default=False), 'hidden\_states\_cache\_dir': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'model\_config\_overrides': FieldInfo(annotation=Dict[str, Any], required=False, default\_factory=dict), 'num\_calibration\_batches': FieldInfo(annotation=int, required=False, default=128), 'num\_draft\_layers': FieldInfo(annotation=int, required=False, default=1), 'sequence\_length': FieldInfo(annotation=int, required=False, default=2048), 'total\_draft\_tokens': FieldInfo(annotation=int, required=False, default=31), 'trimmed\_vocabulary\_length': FieldInfo(annotation=Union[int, NoneType], required=False, default=None), 'trimmed\_vocabulary\_path': FieldInfo(annotation=Union[str, NoneType], required=False, default=None)}*
- Metadata about the fields defined on the model,
mapping of field names to [FieldInfo][pydantic.fields.FieldInfo].
This replaces Model.__fields__ from Pydantic V1.
- *field* num\_calibration\_batches*: int* *= 128*
-
- *field* num\_draft\_layers*: int* *= 1*
-
- *field* sequence\_length*: int* *= 2048*
-
- *field* total\_draft\_tokens*: int* *= 31*
-
- *field* trimmed\_vocabulary\_length*: Optional[int]* *= None*
-
- *field* trimmed\_vocabulary\_path*: Optional[str]* *= None*
-
## EagletV2Qwen35Config
- *class* qairt.experimental.pipeline.torch.llm.eaglet.config.EagletV2Qwen35Config(*\**, *eaglet\_version: str = 'v2\_Qwen35'*, *draft\_safetensors\_path: Optional[str] = None*, *num\_draft\_layers: int = 1*, *model\_config\_overrides: Dict[str, Any] = None*, *num\_calibration\_batches: int = 128*, *calibration\_dataset: str = 'wikitext'*, *sequence\_length: int = 2048*, *dual\_alignment: bool = True*, *encoding\_tie\_recipe\_overrides: Optional[Dict[str, str]] = None*, *dm\_quantsim\_config\_file: Optional[str] = None*, *dsp\_arch: Optional[str] = None*, *dm\_filename\_prefix: str = 'eaglet\_dm'*, *export\_inputs\_embeds: bool = True*, *capture\_hidden\_states: bool = True*, *hidden\_states\_cache\_dir: Optional[str] = None*, *trimmed\_vocabulary\_length: Optional[int] = None*, *trimmed\_vocabulary\_path: Optional[str] = None*, *draft\_top\_k: int = 8*, *draft\_depth: int = 5*, *total\_draft\_tokens: int = 31*, *draft\_preselect: Optional[int] = 64*, *draft\_temperature: Optional[float] = 0.8*, *use\_state\_replay: bool = True*, *gdn\_chunk\_size: int = 128*, *default\_draft\_shape: Literal['tree', 'sequence'] = 'tree'*, *fallback\_to\_sequence\_on\_gdn: bool = False*, *verification\_mode: Literal['auto', 'attention', 'gdn'] = 'auto'*, *\*\*extra\_data: Any*)
- Bases: `EagletV2Config`
Extended configuration for Eaglet v2\_Qwen35 (GDN/linear-attention TM support).
All v2\_Qwen35 additions are TM-side and generator-side. The DM is architecturally
unchanged from v2 — all DM quantization abstractions work without modification.
- eaglet\_version
- Fixed to “v2\_Qwen35” for this variant.
- Type
- str
- use\_state\_replay
- Enable GDN state checkpoint/replay during verification.
- Type
- bool
- gdn\_chunk\_size
- Chunk size for GDN chunkwise-parallel processing.
- Type
- int
- default\_draft\_shape
- Default draft token shape for verification.
“tree” enables tree-structured drafting; “sequence” uses linear chain.
- Type
- Literal[‘tree’, ‘sequence’]
- fallback\_to\_sequence\_on\_gdn
- If True, fall back to sequence verification
for GDN layers instead of tree-masked verification.
- Type
- bool
- verification\_mode
- Verification dispatch strategy.
“auto” selects based on layer attention kinds.
- Type
- Literal[‘auto’, ‘attention’, ‘gdn’]
- export\_inputs\_embeds
- Defaults to True for the Qwen3.5 flow because the
builder supplies the embedding lookup outside the draft graph.
- Type
- bool
- draft\_preselect
- Defaults to 64 (two-stage selection matching on-device behavior).
- Type
- Optional[int]
- draft\_temperature
- Softmax temperature for that two-stage selection.
- Type
- Optional[float]
- *field* default\_draft\_shape*: Literal['tree', 'sequence']* *= 'tree'*
-
- *field* draft\_preselect*: Optional[int]* *= 64*
-
- *field* draft\_temperature*: Optional[float]* *= 0.8*
-
- *field* eaglet\_version*: str* *= 'v2\_Qwen35'*
-
- *field* export\_inputs\_embeds*: bool* *= True*
-
- *field* fallback\_to\_sequence\_on\_gdn*: bool* *= False*
-
- *field* gdn\_chunk\_size*: int* *= 128*
-
- model\_computed\_fields*: ClassVar[dict[str, ComputedFieldInfo]]* *= {}*
- A dictionary of computed field names and their corresponding ComputedFieldInfo objects.
- model\_config*: ClassVar[ConfigDict]* *= {'extra': 'allow', 'protected\_namespaces': ()}*
- Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict].
- model\_fields*: ClassVar[dict[str, FieldInfo]]* *= {'calibration\_dataset': FieldInfo(annotation=str, required=False, default='wikitext'), 'capture\_hidden\_states': FieldInfo(annotation=bool, required=False, default=True), 'default\_draft\_shape': FieldInfo(annotation=Literal['tree', 'sequence'], required=False, default='tree'), 'dm\_filename\_prefix': FieldInfo(annotation=str, required=False, default='eaglet\_dm'), 'dm\_quantsim\_config\_file': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'draft\_depth': FieldInfo(annotation=int, required=False, default=5), 'draft\_preselect': FieldInfo(annotation=Union[int, NoneType], required=False, default=64), 'draft\_safetensors\_path': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'draft\_temperature': FieldInfo(annotation=Union[float, NoneType], required=False, default=0.8), 'draft\_top\_k': FieldInfo(annotation=int, required=False, default=8), 'dsp\_arch': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'dual\_alignment': FieldInfo(annotation=bool, required=False, default=True), 'eaglet\_version': FieldInfo(annotation=str, required=False, default='v2\_Qwen35'), 'encoding\_tie\_recipe\_overrides': FieldInfo(annotation=Union[Dict[str, str], NoneType], required=False, default=None), 'export\_inputs\_embeds': FieldInfo(annotation=bool, required=False, default=True), 'fallback\_to\_sequence\_on\_gdn': FieldInfo(annotation=bool, required=False, default=False), 'gdn\_chunk\_size': FieldInfo(annotation=int, required=False, default=128), 'hidden\_states\_cache\_dir': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'model\_config\_overrides': FieldInfo(annotation=Dict[str, Any], required=False, default\_factory=dict), 'num\_calibration\_batches': FieldInfo(annotation=int, required=False, default=128), 'num\_draft\_layers': FieldInfo(annotation=int, required=False, default=1), 'sequence\_length': FieldInfo(annotation=int, required=False, default=2048), 'total\_draft\_tokens': FieldInfo(annotation=int, required=False, default=31), 'trimmed\_vocabulary\_length': FieldInfo(annotation=Union[int, NoneType], required=False, default=None), 'trimmed\_vocabulary\_path': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'use\_state\_replay': FieldInfo(annotation=bool, required=False, default=True), 'verification\_mode': FieldInfo(annotation=Literal['auto', 'attention', 'gdn'], required=False, default='auto')}*
- Metadata about the fields defined on the model,
mapping of field names to [FieldInfo][pydantic.fields.FieldInfo].
This replaces Model.__fields__ from Pydantic V1.
- *field* use\_state\_replay*: bool* *= True*
-
- *field* verification\_mode*: Literal['auto', 'attention', 'gdn']* *= 'auto'*
-
Last Published: Sep 17, 2026
[Previous Topic
register\_stage\_plugin()](https://docs.qualcomm.com/bundle/publicresource/80-87189-2/topics/qairt-pipeline-base.md)