# LLM Pipeline LLM-specific pipeline implementation, configuration, and context classes. ## LLMPipeline - *class* qairt.experimental.pipeline.torch.llm.pipeline.LLMPipeline(*config: [LLMPipelineConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.LLMPipelineConfig)*) - Bases: [`Pipeline`](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-base.html#qairt.experimental.pipeline.torch.common.bases.pipeline.Pipeline)[[`LLMPipelineConfig`](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.LLMPipelineConfig)] Pipeline for PyTorch LLMs. - *property* config*: [LLMPipelineConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.LLMPipelineConfig)* - Return the domain-specific pipeline configuration. - evaluate(*\*\*kwargs*) → dict[str, float] - Run the configured evaluation metrics on the last stage’s output. Delegates to the last stage’s `evaluate()`. A caller may pass `evaluator_config` via `**kwargs` to override the pipeline-level config. - Returns - `{display_name: score}` e.g. `{"PPL_wikitext": 8.41}`. - Raises - - **RuntimeError** – If `construct()` has not been called, or the last stage produces no evaluable model. - **ValueError** – If no `evaluator_config` or metrics are configured. - *classmethod* from\_pretrained(*model\_id\_or\_path: str*, *recipe: Optional[Union[str, Path, dict[str, Any]]] = None*, *\*\*kwargs: Any*) → [LLMPipeline](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.LLMPipeline) - Create a pipeline from a recipe and execute the `model_loader` stage. - Parameters - - **model\_id\_or\_path** – HuggingFace model ID or local path. - **recipe** – Recipe file path or dict. - **\*\*kwargs** – Forwarded to the base `from_pretrained` as config overrides. - Returns - An `LLMPipeline` instance with the first stage executed. - generate(*prompt: Union[str, list[dict[str, str]], Path]*, *device: Optional[Any] = None*, *\*\*kwargs*) → TextGenerationResult - Generate text using the last stage in the pipeline. The last stage must have `can_generate = True` and implement `generate()`. Users should be aware of the generation capabilities of the last stage and provide appropriate kwargs (e.g. `max_length`, `num_beams`). - Parameters - - **prompt** – Input prompt for generation. Can be a plain string, a list of chat messages (dicts with “role” and “content” keys), or a Path to a JSON file containing chat messages. - **device** – Optional device specification (forwarded to the stage). - **\*\*kwargs** – Additional generation parameters forwarded to the last stage. - Returns - Generation result from the last stage. Stages before `genai_builder` (e.g. `model_loader`, `quantization`) return plain decoded text, which is wrapped in a `TextGenerationResult` for a consistent return type. - Raises - - **RuntimeError** – If `construct()` has not been called or the last stage has not been executed. - **NotImplementedError** – If the last stage does not support generation. ## LLMPipelineConfig - *class* qairt.experimental.pipeline.torch.llm.pipeline.LLMPipelineConfig(*\**, *model\_id\_or\_path: str*, *cache\_dir: str = './workspace'*, *backend: [BackendType](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-api-configs.html#qairt.api.configs.common.BackendType) = BackendType.HTP*, *soc\_details: Optional[Union[qti.aisw.tools.core.utilities.devices.api.device\_definitions.SocDetails, str]] = None*, *log\_level: Optional[str] = None*, *task: str = 'text-generation'*, *features: [PipelineFeatures](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.PipelineFeatures) = None*, *enable\_cache: bool = False*, *checkpoint: Optional[str] = None*, *enable\_observers: bool = False*, *observers: dict[str, Any] = None*, *stage\_info: OrderedDict[str, dict[str, Any]] = None*, *custom\_stage: list[typing.Annotated[pathlib.Path, PathType(path\_type='file')]] = None*, *generator\_config: [qairt.experimental.pipeline.torch.common.configs.GeneratorConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.GeneratorConfig) | None = None*, *evaluator\_config: [qairt.experimental.pipeline.torch.common.configs.EvaluatorConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.EvaluatorConfig) | None = None*, *exporter\_config: [qairt.experimental.pipeline.torch.common.configs.ExporterConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.ExporterConfig) | None = None*, *eaglet: dict[str, Any] | None = None*) - Bases: [`PipelineConfig`](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-base.html#qairt.experimental.pipeline.torch.common.bases.pipeline.PipelineConfig), [`LLMPipelineContext`](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.LLMPipelineContext) Global pipeline configuration for LLM pipelines. - add\_config(*stage\_name: str*, *config: [qairt.experimental.pipeline.torch.common.configs.ExporterConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.ExporterConfig) | [qairt.experimental.pipeline.torch.common.configs.GeneratorConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.GeneratorConfig) | [qairt.experimental.pipeline.torch.common.configs.EvaluatorConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.EvaluatorConfig)*) → None - Adds a config for a specific stage. - Parameters - - **stage\_name** – Name of the stage to associate the config with. - **config** – Must be one of ExporterConfig, GeneratorConfig or EvaluatorConfig. Stage-specific configs override global configs for the same stage. - add\_dataloader(*stage\_name: str*, *dataloader: DataLoader*) → None - Adds a data loader for a specific stage. - Parameters - - **stage\_name** – Name of the stage to associate the dataloader with (e.g., “quantization” for calibration dataloader) - **dataloader** – The DataLoader instance to add - *field* eaglet*: dict[str, Any] | None* *= None* - - Validated by - - `_check_unsupported_features` - *field* evaluator\_config*: [EvaluatorConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.EvaluatorConfig) | None* *= None* - - Validated by - - `_check_unsupported_features` - *field* exporter\_config*: [ExporterConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.ExporterConfig) | None* *= None* - - Validated by - - `_check_unsupported_features` - *classmethod* from\_recipe(*recipe: Union[str, Path, dict[str, Any]]*) → [LLMPipelineConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.LLMPipelineConfig) - Load pipeline configuration from a YAML recipe file or dict. - Parameters - **recipe** – Path to YAML recipe file, or a pre-loaded recipe dict. - Returns - `LLMPipelineConfig` instance. - *field* generator\_config*: [GeneratorConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-common.html#qairt.experimental.pipeline.torch.common.configs.GeneratorConfig) | None* *= None* - - Validated by - - `_check_unsupported_features` - model\_computed\_fields*: ClassVar[dict[str, ComputedFieldInfo]]* *= {}* - A dictionary of computed field names and their corresponding ComputedFieldInfo objects. - model\_config*: ClassVar[ConfigDict]* *= {'arbitrary\_types\_allowed': True, 'protected\_namespaces': ()}* - Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict]. - model\_fields*: ClassVar[dict[str, FieldInfo]]* *= {'backend': FieldInfo(annotation=BackendType, required=False, default=<BackendType.HTP: 'HTP'>), 'cache\_dir': FieldInfo(annotation=str, required=False, default='./workspace'), 'checkpoint': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'custom\_stage': FieldInfo(annotation=list[Annotated[Path, PathType]], required=False, default\_factory=list), 'eaglet': FieldInfo(annotation=Union[dict[str, Any], NoneType], required=False, default=None), 'enable\_cache': FieldInfo(annotation=bool, required=False, default=False), 'enable\_observers': FieldInfo(annotation=bool, required=False, default=False), 'evaluator\_config': FieldInfo(annotation=Union[EvaluatorConfig, NoneType], required=False, default=None), 'exporter\_config': FieldInfo(annotation=Union[ExporterConfig, NoneType], required=False, default=None), 'features': FieldInfo(annotation=PipelineFeatures, required=False, default\_factory=PipelineFeatures), 'generator\_config': FieldInfo(annotation=Union[GeneratorConfig, NoneType], required=False, default=None), 'log\_level': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'model\_id\_or\_path': FieldInfo(annotation=str, required=True), 'observers': FieldInfo(annotation=dict[str, Any], required=False, default\_factory=dict), 'soc\_details': FieldInfo(annotation=Union[SocDetails, str, NoneType], required=False, default=None), 'stage\_info': FieldInfo(annotation=OrderedDict[str, dict[str, Any]], required=False, default\_factory=OrderedDict, alias\_priority=2, serialization\_alias='stages'), 'task': FieldInfo(annotation=str, required=False, default='text-generation')}* - Metadata about the fields defined on the model, mapping of field names to [FieldInfo][pydantic.fields.FieldInfo]. This replaces Model.__fields__ from Pydantic V1. - model\_post\_init(*context: Any*, */*) → None - We need to both initialize private attributes and call the user-defined model\_post\_init method. ## LLMPipelineContext - *class* qairt.experimental.pipeline.torch.llm.pipeline.LLMPipelineContext(*\**, *model\_id\_or\_path: str*, *cache\_dir: str = './workspace'*, *backend: [BackendType](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-api-configs.html#qairt.api.configs.common.BackendType) = BackendType.HTP*, *soc\_details: Optional[Union[qti.aisw.tools.core.utilities.devices.api.device\_definitions.SocDetails, str]] = None*, *log\_level: Optional[str] = None*, *task: str = 'text-generation'*, *features: [PipelineFeatures](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.PipelineFeatures) = None*) - Bases: [`PipelineContext`](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-base.html#qairt.experimental.pipeline.torch.common.bases.pipeline_context.PipelineContext) LLM-specific stage-visible pipeline context. Extends `PipelineContext` with fields specific to large language models. Injected into `StageConfig._pipeline_context` for all LLM stages. - *field* features*: [PipelineFeatures](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-llm.html#qairt.experimental.pipeline.torch.llm.pipeline.PipelineFeatures)* *[Optional]* - - model\_computed\_fields*: ClassVar[dict[str, ComputedFieldInfo]]* *= {}* - A dictionary of computed field names and their corresponding ComputedFieldInfo objects. - model\_config*: ClassVar[ConfigDict]* *= {'arbitrary\_types\_allowed': True, 'protected\_namespaces': ()}* - Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict]. - model\_fields*: ClassVar[dict[str, FieldInfo]]* *= {'backend': FieldInfo(annotation=BackendType, required=False, default=<BackendType.HTP: 'HTP'>), 'cache\_dir': FieldInfo(annotation=str, required=False, default='./workspace'), 'features': FieldInfo(annotation=PipelineFeatures, required=False, default\_factory=PipelineFeatures), 'log\_level': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'model\_id\_or\_path': FieldInfo(annotation=str, required=True), 'soc\_details': FieldInfo(annotation=Union[SocDetails, str, NoneType], required=False, default=None), 'task': FieldInfo(annotation=str, required=False, default='text-generation')}* - Metadata about the fields defined on the model, mapping of field names to [FieldInfo][pydantic.fields.FieldInfo]. This replaces Model.__fields__ from Pydantic V1. - model\_post\_init(*context: Any*, */*) → None - We need to both initialize private attributes and call the user-defined model\_post\_init method. - *field* task*: str* *= 'text-generation'* - ## PipelineFeatures - *class* qairt.experimental.pipeline.torch.llm.pipeline.PipelineFeatures(*\**, *speculative\_decoding: Tuple[bool, str | None] = (False, None)*, *lora: Optional[[LoRAFeatureConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-lora.html#qairt.experimental.pipeline.torch.llm.lora.configs.LoRAFeatureConfig)] = None*) - Bases: `BaseModel` Cross-cutting features that can be enabled/disabled for the entire pipeline. These features can affect multiple stages. If set: - stages will check to ensure feature related configs are provided. - stages will enable feature-specific logic during execution. **Field declaration order is load-bearing**: GeneratorFactory composes generator mixins in the order fields are declared here. The first-declared field gets the highest MRO position and therefore runs first in forward(). Do not reorder fields without updating the MRO contract test. - feature\_keys() → list[str] - Derive all active feature keys in field declaration order. Iterates over the model fields in the order they are declared and collects a key for every enabled field. For tuple-valued fields (e.g. `speculative_decoding = (True, "eaglet")`) the key is `"{field_name}_{method}"`; plain `bool` or `Optional` fields use the field name directly. **The returned order is load-bearing**: it determines mixin MRO position inside `GeneratorFactory._compose()`. The first key in the list produces the highest mixin in the MRO (called first in `forward()`). Reordering fields in `PipelineFeatures` therefore changes execution order — do not do so without updating the MRO contract test. - Returns - Ordered list of active feature key strings, in field declaration order. Empty when no feature is enabled. - *field* lora*: Optional[[LoRAFeatureConfig](https://docs.qualcomm.com/doc/80-87189-2/topic/qairt-pipeline-lora.html#qairt.experimental.pipeline.torch.llm.lora.configs.LoRAFeatureConfig)]* *= None* - - model\_computed\_fields*: ClassVar[dict[str, ComputedFieldInfo]]* *= {}* - A dictionary of computed field names and their corresponding ComputedFieldInfo objects. - model\_config*: ClassVar[ConfigDict]* *= {}* - Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict]. - model\_fields*: ClassVar[dict[str, FieldInfo]]* *= {'lora': FieldInfo(annotation=Union[LoRAFeatureConfig, NoneType], required=False, default=None), 'speculative\_decoding': FieldInfo(annotation=Tuple[bool, Union[str, NoneType]], required=False, default=(False, None))}* - Metadata about the fields defined on the model, mapping of field names to [FieldInfo][pydantic.fields.FieldInfo]. This replaces Model.__fields__ from Pydantic V1. - *field* speculative\_decoding*: Tuple[bool, str | None]* *= (False, None)* - ## EagletV2Config - *class* qairt.experimental.pipeline.torch.llm.eaglet.config.EagletV2Config(*\**, *eaglet\_version: str = 'v2'*, *draft\_safetensors\_path: Optional[str] = None*, *num\_draft\_layers: int = 1*, *model\_config\_overrides: Dict[str, Any] = None*, *num\_calibration\_batches: int = 128*, *calibration\_dataset: str = 'wikitext'*, *sequence\_length: int = 2048*, *dual\_alignment: bool = True*, *encoding\_tie\_recipe\_overrides: Optional[Dict[str, str]] = None*, *dm\_quantsim\_config\_file: Optional[str] = None*, *dsp\_arch: Optional[str] = None*, *dm\_filename\_prefix: str = 'eaglet\_dm'*, *export\_inputs\_embeds: bool = False*, *capture\_hidden\_states: bool = True*, *hidden\_states\_cache\_dir: Optional[str] = None*, *trimmed\_vocabulary\_length: Optional[int] = None*, *trimmed\_vocabulary\_path: Optional[str] = None*, *draft\_top\_k: int = 8*, *draft\_depth: int = 5*, *total\_draft\_tokens: int = 31*, *draft\_preselect: Optional[int] = None*, *draft\_temperature: Optional[float] = None*, *\*\*extra\_data: Any*) - Bases: `BaseModel` Configuration for Eaglet v2 speculative decoding pipeline. Controls draft model construction, calibration, and encoding-tie behavior. Used by the dm\_model\_loading and dm\_quantization stages. The draft spec family is resolved automatically from the TM’s `config.model_type` via `MODEL_TYPE_TO_FAMILY`. Users do not need to specify it. - eaglet\_version - Eaglet version string (e.g. “v2”, “v2\_Qwen35”). - Type - str - draft\_safetensors\_path - Path to pretrained DM weights. None = random init. - Type - Optional[str] - num\_draft\_layers - Number of eaglet decoder layers in the draft model. - Type - int - num\_calibration\_batches - Number of batches for DM calibration. - Type - int - calibration\_dataset - Dataset name for calibration data. - Type - str - sequence\_length - Calibration dataloader context length (sequence length). - Type - int - dual\_alignment - Whether to run cross\_model\_tie for encoding alignment. - Type - bool - encoding\_tie\_recipe\_overrides - Override default ONNX op names for encoding tie. Keys must match EncodingTieRecipe field names. - Type - Optional[Dict[str, str]] - dm\_quantsim\_config\_file - Path to HTP quantsim config JSON for per-channel weight quantization. None = use default AIMET config. - Type - Optional[str] - dsp\_arch - DSP architecture version for HTP config selection (e.g. “v79”, “v81”). None = use pipeline recipe value or CalibratorParams default. - Type - Optional[str] - dm\_filename\_prefix - Prefix for exported DM artifact files. - Type - str - export\_inputs\_embeds - Export the DM with `inputs_embeds` instead of `input_ids` plus an in-graph embedding lookup. - Type - bool - capture\_hidden\_states - Whether TM quantization should capture hidden states for DM calibration. - Type - bool - hidden\_states\_cache\_dir - Directory to store captured TM hidden states. None = use pipeline cache\_dir. - Type - Optional[str] - draft\_top\_k - Number of top-k candidates during tree drafting. - Type - int - draft\_depth - Maximum tree depth for speculative drafting. - Type - int - total\_draft\_tokens - Maximum number of draft tokens per step. - Type - int - draft\_preselect - Preselect count for the two-stage token selection (top-k → softmax → top-k). None uses log-softmax → top-k. - Type - Optional[int] - draft\_temperature - Softmax temperature for the two-stage selection. Ignored when `draft_preselect` is None. - Type - Optional[float] - model\_config\_overrides - DM HF config overrides keyed by attribute name (e.g. `{"hidden_size": 1024}`). Takes priority over the pretrained DM config.json and the TM value. Unknown keys are set verbatim. - Type - Dict[str, Any] - *field* calibration\_dataset*: str* *= 'wikitext'* - - *field* capture\_hidden\_states*: bool* *= True* - - *field* dm\_filename\_prefix*: str* *= 'eaglet\_dm'* - - *field* dm\_quantsim\_config\_file*: Optional[str]* *= None* - - *field* draft\_depth*: int* *= 5* - - *field* draft\_preselect*: Optional[int]* *= None* - - *field* draft\_safetensors\_path*: Optional[str]* *= None* - - *field* draft\_temperature*: Optional[float]* *= None* - - *field* draft\_top\_k*: int* *= 8* - - *field* dsp\_arch*: Optional[str]* *= None* - - *field* dual\_alignment*: bool* *= True* - - *field* eaglet\_version*: str* *= 'v2'* - - *field* encoding\_tie\_recipe\_overrides*: Optional[Dict[str, str]]* *= None* - - *field* export\_inputs\_embeds*: bool* *= False* - - *field* hidden\_states\_cache\_dir*: Optional[str]* *= None* - - model\_computed\_fields*: ClassVar[dict[str, ComputedFieldInfo]]* *= {}* - A dictionary of computed field names and their corresponding ComputedFieldInfo objects. - model\_config*: ClassVar[ConfigDict]* *= {'extra': 'allow', 'protected\_namespaces': ()}* - Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict]. - *field* model\_config\_overrides*: Dict[str, Any]* *[Optional]* - - model\_fields*: ClassVar[dict[str, FieldInfo]]* *= {'calibration\_dataset': FieldInfo(annotation=str, required=False, default='wikitext'), 'capture\_hidden\_states': FieldInfo(annotation=bool, required=False, default=True), 'dm\_filename\_prefix': FieldInfo(annotation=str, required=False, default='eaglet\_dm'), 'dm\_quantsim\_config\_file': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'draft\_depth': FieldInfo(annotation=int, required=False, default=5), 'draft\_preselect': FieldInfo(annotation=Union[int, NoneType], required=False, default=None), 'draft\_safetensors\_path': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'draft\_temperature': FieldInfo(annotation=Union[float, NoneType], required=False, default=None), 'draft\_top\_k': FieldInfo(annotation=int, required=False, default=8), 'dsp\_arch': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'dual\_alignment': FieldInfo(annotation=bool, required=False, default=True), 'eaglet\_version': FieldInfo(annotation=str, required=False, default='v2'), 'encoding\_tie\_recipe\_overrides': FieldInfo(annotation=Union[Dict[str, str], NoneType], required=False, default=None), 'export\_inputs\_embeds': FieldInfo(annotation=bool, required=False, default=False), 'hidden\_states\_cache\_dir': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'model\_config\_overrides': FieldInfo(annotation=Dict[str, Any], required=False, default\_factory=dict), 'num\_calibration\_batches': FieldInfo(annotation=int, required=False, default=128), 'num\_draft\_layers': FieldInfo(annotation=int, required=False, default=1), 'sequence\_length': FieldInfo(annotation=int, required=False, default=2048), 'total\_draft\_tokens': FieldInfo(annotation=int, required=False, default=31), 'trimmed\_vocabulary\_length': FieldInfo(annotation=Union[int, NoneType], required=False, default=None), 'trimmed\_vocabulary\_path': FieldInfo(annotation=Union[str, NoneType], required=False, default=None)}* - Metadata about the fields defined on the model, mapping of field names to [FieldInfo][pydantic.fields.FieldInfo]. This replaces Model.__fields__ from Pydantic V1. - *field* num\_calibration\_batches*: int* *= 128* - - *field* num\_draft\_layers*: int* *= 1* - - *field* sequence\_length*: int* *= 2048* - - *field* total\_draft\_tokens*: int* *= 31* - - *field* trimmed\_vocabulary\_length*: Optional[int]* *= None* - - *field* trimmed\_vocabulary\_path*: Optional[str]* *= None* - ## EagletV2Qwen35Config - *class* qairt.experimental.pipeline.torch.llm.eaglet.config.EagletV2Qwen35Config(*\**, *eaglet\_version: str = 'v2\_Qwen35'*, *draft\_safetensors\_path: Optional[str] = None*, *num\_draft\_layers: int = 1*, *model\_config\_overrides: Dict[str, Any] = None*, *num\_calibration\_batches: int = 128*, *calibration\_dataset: str = 'wikitext'*, *sequence\_length: int = 2048*, *dual\_alignment: bool = True*, *encoding\_tie\_recipe\_overrides: Optional[Dict[str, str]] = None*, *dm\_quantsim\_config\_file: Optional[str] = None*, *dsp\_arch: Optional[str] = None*, *dm\_filename\_prefix: str = 'eaglet\_dm'*, *export\_inputs\_embeds: bool = True*, *capture\_hidden\_states: bool = True*, *hidden\_states\_cache\_dir: Optional[str] = None*, *trimmed\_vocabulary\_length: Optional[int] = None*, *trimmed\_vocabulary\_path: Optional[str] = None*, *draft\_top\_k: int = 8*, *draft\_depth: int = 5*, *total\_draft\_tokens: int = 31*, *draft\_preselect: Optional[int] = 64*, *draft\_temperature: Optional[float] = 0.8*, *use\_state\_replay: bool = True*, *gdn\_chunk\_size: int = 128*, *default\_draft\_shape: Literal['tree', 'sequence'] = 'tree'*, *fallback\_to\_sequence\_on\_gdn: bool = False*, *verification\_mode: Literal['auto', 'attention', 'gdn'] = 'auto'*, *\*\*extra\_data: Any*) - Bases: `EagletV2Config` Extended configuration for Eaglet v2\_Qwen35 (GDN/linear-attention TM support). All v2\_Qwen35 additions are TM-side and generator-side. The DM is architecturally unchanged from v2 — all DM quantization abstractions work without modification. - eaglet\_version - Fixed to “v2\_Qwen35” for this variant. - Type - str - use\_state\_replay - Enable GDN state checkpoint/replay during verification. - Type - bool - gdn\_chunk\_size - Chunk size for GDN chunkwise-parallel processing. - Type - int - default\_draft\_shape - Default draft token shape for verification. “tree” enables tree-structured drafting; “sequence” uses linear chain. - Type - Literal[‘tree’, ‘sequence’] - fallback\_to\_sequence\_on\_gdn - If True, fall back to sequence verification for GDN layers instead of tree-masked verification. - Type - bool - verification\_mode - Verification dispatch strategy. “auto” selects based on layer attention kinds. - Type - Literal[‘auto’, ‘attention’, ‘gdn’] - export\_inputs\_embeds - Defaults to True for the Qwen3.5 flow because the builder supplies the embedding lookup outside the draft graph. - Type - bool - draft\_preselect - Defaults to 64 (two-stage selection matching on-device behavior). - Type - Optional[int] - draft\_temperature - Softmax temperature for that two-stage selection. - Type - Optional[float] - *field* default\_draft\_shape*: Literal['tree', 'sequence']* *= 'tree'* - - *field* draft\_preselect*: Optional[int]* *= 64* - - *field* draft\_temperature*: Optional[float]* *= 0.8* - - *field* eaglet\_version*: str* *= 'v2\_Qwen35'* - - *field* export\_inputs\_embeds*: bool* *= True* - - *field* fallback\_to\_sequence\_on\_gdn*: bool* *= False* - - *field* gdn\_chunk\_size*: int* *= 128* - - model\_computed\_fields*: ClassVar[dict[str, ComputedFieldInfo]]* *= {}* - A dictionary of computed field names and their corresponding ComputedFieldInfo objects. - model\_config*: ClassVar[ConfigDict]* *= {'extra': 'allow', 'protected\_namespaces': ()}* - Configuration for the model, should be a dictionary conforming to [ConfigDict][pydantic.config.ConfigDict]. - model\_fields*: ClassVar[dict[str, FieldInfo]]* *= {'calibration\_dataset': FieldInfo(annotation=str, required=False, default='wikitext'), 'capture\_hidden\_states': FieldInfo(annotation=bool, required=False, default=True), 'default\_draft\_shape': FieldInfo(annotation=Literal['tree', 'sequence'], required=False, default='tree'), 'dm\_filename\_prefix': FieldInfo(annotation=str, required=False, default='eaglet\_dm'), 'dm\_quantsim\_config\_file': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'draft\_depth': FieldInfo(annotation=int, required=False, default=5), 'draft\_preselect': FieldInfo(annotation=Union[int, NoneType], required=False, default=64), 'draft\_safetensors\_path': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'draft\_temperature': FieldInfo(annotation=Union[float, NoneType], required=False, default=0.8), 'draft\_top\_k': FieldInfo(annotation=int, required=False, default=8), 'dsp\_arch': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'dual\_alignment': FieldInfo(annotation=bool, required=False, default=True), 'eaglet\_version': FieldInfo(annotation=str, required=False, default='v2\_Qwen35'), 'encoding\_tie\_recipe\_overrides': FieldInfo(annotation=Union[Dict[str, str], NoneType], required=False, default=None), 'export\_inputs\_embeds': FieldInfo(annotation=bool, required=False, default=True), 'fallback\_to\_sequence\_on\_gdn': FieldInfo(annotation=bool, required=False, default=False), 'gdn\_chunk\_size': FieldInfo(annotation=int, required=False, default=128), 'hidden\_states\_cache\_dir': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'model\_config\_overrides': FieldInfo(annotation=Dict[str, Any], required=False, default\_factory=dict), 'num\_calibration\_batches': FieldInfo(annotation=int, required=False, default=128), 'num\_draft\_layers': FieldInfo(annotation=int, required=False, default=1), 'sequence\_length': FieldInfo(annotation=int, required=False, default=2048), 'total\_draft\_tokens': FieldInfo(annotation=int, required=False, default=31), 'trimmed\_vocabulary\_length': FieldInfo(annotation=Union[int, NoneType], required=False, default=None), 'trimmed\_vocabulary\_path': FieldInfo(annotation=Union[str, NoneType], required=False, default=None), 'use\_state\_replay': FieldInfo(annotation=bool, required=False, default=True), 'verification\_mode': FieldInfo(annotation=Literal['auto', 'attention', 'gdn'], required=False, default='auto')}* - Metadata about the fields defined on the model, mapping of field names to [FieldInfo][pydantic.fields.FieldInfo]. This replaces Model.__fields__ from Pydantic V1. - *field* use\_state\_replay*: bool* *= True* - - *field* verification\_mode*: Literal['auto', 'attention', 'gdn']* *= 'auto'* - Last Published: Sep 17, 2026 [Previous Topic register\_stage\_plugin()](https://docs.qualcomm.com/bundle/publicresource/80-87189-2/topics/qairt-pipeline-base.md)