Index Symbols | _ | A | B | C | D | E | F | G | H | I | J | K | L | M | N | O | P | Q | R | S | T | U | V | W | Y | Z Symbols --accuracy_threshold trtllm-eval-mmlu command line option --agent_percentage trtllm-serve-serve command line option --agent_types trtllm-serve-serve command line option --apply_chat_template trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-cnn_dailymail command line option trtllm-eval-covost2 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-longbench_v1 command line option trtllm-eval-longbench_v2 command line option trtllm-eval-mmlu command line option --backend trtllm-bench-latency command line option trtllm-bench-throughput command line option trtllm-eval command line option trtllm-serve-serve command line option --beam_width trtllm-bench-latency command line option trtllm-bench-throughput command line option --chat_template trtllm-serve-serve command line option --chat_template_kwargs trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-longbench_v1 command line option trtllm-eval-longbench_v2 command line option trtllm-eval-mmlu command line option trtllm-eval-mmmu command line option trtllm-eval-mmmu_pro command line option --check_accuracy trtllm-eval-mmlu command line option --cluster_size trtllm-bench-throughput command line option trtllm-serve-serve command line option --concurrency trtllm-bench-latency command line option trtllm-bench-throughput command line option --config trtllm-bench-latency command line option trtllm-bench-throughput command line option trtllm-eval command line option trtllm-serve-disaggregated command line option trtllm-serve-disaggregated_mpi_worker command line option trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --config_file trtllm-serve-disaggregated command line option trtllm-serve-disaggregated_mpi_worker command line option --context_parallel_size trtllm-serve-serve command line option --cot trtllm-eval-longbench_v2 command line option --cp_size trtllm-serve-serve command line option --custom_module_dirs trtllm-bench-throughput command line option trtllm-serve-serve command line option --custom_tokenizer trtllm-bench-latency command line option trtllm-bench-throughput command line option trtllm-eval command line option trtllm-serve-serve command line option --data_device trtllm-bench-throughput command line option --dataset trtllm-bench-build command line option trtllm-bench-latency command line option trtllm-bench-throughput command line option --dataset_path trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-cnn_dailymail command line option trtllm-eval-covost2 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-json_mode_eval command line option trtllm-eval-longbench_v1 command line option trtllm-eval-longbench_v2 command line option trtllm-eval-mmlu command line option trtllm-eval-mmmu command line option trtllm-eval-mmmu_pro command line option --difficulty trtllm-eval-longbench_v2 command line option --disable_chunked_context trtllm-bench-throughput command line option --disable_kv_cache_reuse trtllm-eval command line option --disagg_cluster_uri trtllm-serve-serve command line option --domain trtllm-eval-longbench_v2 command line option --dump_samples_path trtllm-eval-covost2 command line option --enable_attention_dp trtllm-serve-serve command line option --enable_chunked_context trtllm-bench-throughput command line option --enable_chunked_prefill trtllm-serve-serve command line option --engine_dir trtllm-bench-latency command line option trtllm-bench-throughput command line option --eos_id trtllm-bench-throughput command line option --ep trtllm-bench-latency command line option trtllm-bench-throughput command line option --ep_size trtllm-eval command line option trtllm-serve-serve command line option --extra_encoder_options trtllm-serve-mm_embedding_serve command line option --extra_llm_api_options trtllm-bench-latency command line option trtllm-bench-throughput command line option trtllm-eval command line option trtllm-serve-serve command line option --extra_visual_gen_options trtllm-serve-serve command line option --fail_fast_on_attention_window_too_large trtllm-serve-serve command line option --fewshot_as_multiturn trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-gsm8k command line option --free_gpu_memory_fraction trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --gpus_per_node trtllm-eval command line option trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --grpc trtllm-serve-serve command line option --hf_revision trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --host trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --image_data_format trtllm-bench-throughput command line option --iteration_log trtllm-bench-latency command line option trtllm-bench-throughput command line option --kv_cache_dtype trtllm-serve-serve command line option --kv_cache_free_gpu_mem_fraction trtllm-bench-latency command line option trtllm-bench-throughput command line option --kv_cache_free_gpu_memory_fraction trtllm-eval command line option trtllm-serve-serve command line option --lang_pair trtllm-eval-covost2 command line option --length trtllm-eval-longbench_v2 command line option --log_level trtllm-bench command line option trtllm-eval command line option trtllm-serve-disaggregated command line option trtllm-serve-disaggregated_mpi_worker command line option trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --log_samples trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-longbench_v1 command line option trtllm-eval-mmmu command line option trtllm-eval-mmmu_pro command line option --max_batch_size trtllm-bench-build command line option trtllm-bench-throughput command line option trtllm-eval command line option trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --max_beam_width trtllm-eval command line option trtllm-serve-serve command line option --max_input_len trtllm-bench-latency command line option trtllm-bench-throughput command line option --max_input_length trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-cnn_dailymail command line option trtllm-eval-covost2 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-json_mode_eval command line option trtllm-eval-longbench_v2 command line option trtllm-eval-mmlu command line option trtllm-eval-mmmu command line option trtllm-eval-mmmu_pro command line option --max_len trtllm-eval-longbench_v2 command line option --max_num_tokens trtllm-bench-build command line option trtllm-bench-throughput command line option trtllm-eval command line option trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --max_output_length trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-cnn_dailymail command line option trtllm-eval-covost2 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-json_mode_eval command line option trtllm-eval-longbench_v2 command line option trtllm-eval-mmlu command line option trtllm-eval-mmmu command line option trtllm-eval-mmmu_pro command line option --max_seq_len trtllm-bench-build command line option trtllm-bench-latency command line option trtllm-bench-throughput command line option trtllm-eval command line option trtllm-serve-serve command line option --media_io_kwargs trtllm-serve-serve command line option --medusa_choices trtllm-bench-latency command line option --metadata_server_config_file trtllm-serve-disaggregated command line option trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --metrics-log-interval trtllm-serve-disaggregated command line option --middleware trtllm-serve-serve command line option --modality trtllm-bench-latency command line option trtllm-bench-throughput command line option --model trtllm-bench command line option trtllm-eval command line option --model_path trtllm-bench command line option --moe_cluster_parallel_size trtllm-serve-serve command line option --moe_expert_parallel_size trtllm-serve-serve command line option --no-telemetry trtllm-bench command line option trtllm-eval command line option trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --no_context trtllm-eval-longbench_v2 command line option --no_skip_tokenizer_init trtllm-bench-throughput command line option --no_weights_loading trtllm-bench-build command line option --num_fewshot trtllm-eval-mmlu command line option --num_postprocess_workers trtllm-serve-serve command line option --num_requests trtllm-bench-latency command line option trtllm-bench-throughput command line option --num_samples trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-cnn_dailymail command line option trtllm-eval-covost2 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-json_mode_eval command line option trtllm-eval-longbench_v1 command line option trtllm-eval-longbench_v2 command line option trtllm-eval-mmlu command line option trtllm-eval-mmmu command line option trtllm-eval-mmmu_pro command line option --otlp_traces_endpoint trtllm-serve-serve command line option --output_dir trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-cnn_dailymail command line option trtllm-eval-covost2 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-json_mode_eval command line option trtllm-eval-longbench_v1 command line option trtllm-eval-longbench_v2 command line option trtllm-eval-mmlu command line option trtllm-eval-mmmu command line option trtllm-eval-mmmu_pro command line option --output_json trtllm-bench-throughput command line option --output_path trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-longbench_v1 command line option trtllm-eval-mmmu command line option trtllm-eval-mmmu_pro command line option --pipeline_parallel_size trtllm-serve-serve command line option --port trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --pp trtllm-bench-latency command line option trtllm-bench-throughput command line option --pp_size trtllm-bench-build command line option trtllm-eval command line option trtllm-serve-serve command line option --prompts_dir trtllm-eval-longbench_v2 command line option --quantization trtllm-bench-build command line option --rag trtllm-eval-longbench_v2 command line option --random_seed trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-cnn_dailymail command line option trtllm-eval-covost2 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-json_mode_eval command line option trtllm-eval-longbench_v1 command line option trtllm-eval-longbench_v2 command line option trtllm-eval-mmlu command line option trtllm-eval-mmmu command line option trtllm-eval-mmmu_pro command line option --reasoning_parser trtllm-serve-serve command line option --report_json trtllm-bench-latency command line option trtllm-bench-throughput command line option --request_json trtllm-bench-throughput command line option --request_timeout trtllm-serve-disaggregated command line option --revision trtllm-bench command line option trtllm-eval command line option trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --rouge_path trtllm-eval-cnn_dailymail command line option --sampler_options trtllm-bench-latency command line option trtllm-bench-throughput command line option --sampling_seed trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-mmmu_pro command line option --schedule_style trtllm-serve-disaggregated command line option --scheduler_policy trtllm-bench-throughput command line option --served_model_name trtllm-serve-serve command line option --server_role trtllm-serve-serve command line option --server_start_timeout trtllm-serve-disaggregated command line option --start_idx trtllm-eval-longbench_v2 command line option --streaming trtllm-bench-throughput command line option --subset trtllm-eval-mmmu_pro command line option --system_prompt trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-cnn_dailymail command line option trtllm-eval-covost2 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-json_mode_eval command line option trtllm-eval-longbench_v1 command line option trtllm-eval-longbench_v2 command line option trtllm-eval-mmlu command line option trtllm-eval-mmmu command line option trtllm-eval-mmmu_pro command line option --target_input_len trtllm-bench-build command line option trtllm-bench-throughput command line option --target_output_len trtllm-bench-build command line option trtllm-bench-throughput command line option --telemetry trtllm-bench command line option trtllm-eval command line option trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --temperature trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-covost2 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-longbench_v2 command line option trtllm-eval-mmmu_pro command line option --tensor_parallel_size trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --tokenizer trtllm-eval command line option trtllm-serve-serve command line option --tool_parser trtllm-serve-serve command line option --top_k trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-mmmu_pro command line option --top_p trtllm-eval-aime25 command line option trtllm-eval-aime26 command line option trtllm-eval-gpqa_diamond command line option trtllm-eval-gpqa_extended command line option trtllm-eval-gpqa_main command line option trtllm-eval-gsm8k command line option trtllm-eval-longbench_v2 command line option trtllm-eval-mmmu_pro command line option --tp trtllm-bench-latency command line option trtllm-bench-throughput command line option --tp_size trtllm-bench-build command line option trtllm-eval command line option trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --trust_remote_code trtllm-bench-build command line option trtllm-eval command line option trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option --video_pruning_rate trtllm-serve-serve command line option --warmup trtllm-bench-latency command line option trtllm-bench-throughput command line option --workspace trtllm-bench command line option -c trtllm-serve-disaggregated command line option trtllm-serve-disaggregated_mpi_worker command line option -l trtllm-serve-disaggregated command line option -m trtllm-bench command line option trtllm-serve-disaggregated command line option -pp trtllm-bench-build command line option -q trtllm-bench-build command line option -r trtllm-serve-disaggregated command line option -s trtllm-serve-disaggregated command line option -t trtllm-serve-disaggregated command line option -tp trtllm-bench-build command line option -w trtllm-bench command line option _ __init__() (tensorrt_llm.llmapi.AsyncLLM method) (tensorrt_llm.llmapi.AttentionDpConfig method) (tensorrt_llm.llmapi.AutoDecodingConfig method) (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.BuildCacheConfig method) (tensorrt_llm.llmapi.BuildConfig method) (tensorrt_llm.llmapi.CacheTransceiverConfig method) (tensorrt_llm.llmapi.CalibConfig method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.CompletionOutput method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.CudaGraphConfig method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig method) (tensorrt_llm.llmapi.DFlashDecodingConfig method) (tensorrt_llm.llmapi.DisaggregatedParams method) (tensorrt_llm.llmapi.DisaggScheduleStyle method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig method) (tensorrt_llm.llmapi.DynamicBatchConfig method) (tensorrt_llm.llmapi.Eagle3DecodingConfig method) (tensorrt_llm.llmapi.EagleDecodingConfig method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig method) (tensorrt_llm.llmapi.GuidedDecodingParams method) (tensorrt_llm.llmapi.KvCacheConfig method) (tensorrt_llm.llmapi.KvCacheRetentionConfig method) (tensorrt_llm.llmapi.KvCacheRetentionConfig.TokenRangeRetentionConfig method) (tensorrt_llm.llmapi.LLM method) (tensorrt_llm.llmapi.LookaheadDecodingConfig method) (tensorrt_llm.llmapi.LoRARequest method) (tensorrt_llm.llmapi.MedusaDecodingConfig method) (tensorrt_llm.llmapi.MoeConfig method) (tensorrt_llm.llmapi.MpiCommSession method) (tensorrt_llm.llmapi.MTPDecodingConfig method) (tensorrt_llm.llmapi.MultimodalEncoder method) (tensorrt_llm.llmapi.NGramDecodingConfig method) (tensorrt_llm.llmapi.PARDDecodingConfig method) (tensorrt_llm.llmapi.PrometheusMetricsConfig method) (tensorrt_llm.llmapi.QuantAlgo method) (tensorrt_llm.llmapi.QuantConfig method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig method) (tensorrt_llm.llmapi.RequestError method) (tensorrt_llm.llmapi.RequestOutput method) (tensorrt_llm.llmapi.RequestOutput.PostprocWorker method) (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Input method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig method) (tensorrt_llm.llmapi.SADecodingConfig method) (tensorrt_llm.llmapi.SAEnhancerConfig method) (tensorrt_llm.llmapi.SamplingParams method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig method) (tensorrt_llm.llmapi.SchedulerConfig method) (tensorrt_llm.llmapi.SchedulingParams method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig method) (tensorrt_llm.llmapi.TorchCompileConfig method) (tensorrt_llm.llmapi.TorchLlmArgs method) (tensorrt_llm.llmapi.TrtLlmArgs method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig method) A abort() (tensorrt_llm.llmapi.MpiCommSession method) (tensorrt_llm.llmapi.RequestOutput method) aborted() (tensorrt_llm.llmapi.RequestOutput method) abs() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) acceptance_length_threshold (tensorrt_llm.llmapi.AutoDecodingConfig attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) acceptance_window (tensorrt_llm.llmapi.AutoDecodingConfig attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) activation() (in module tensorrt_llm.functional) AdaLayerNorm (class in tensorrt_llm.layers.normalization) AdaLayerNormContinuous (class in tensorrt_llm.layers.normalization) AdaLayerNormZero (class in tensorrt_llm.layers.normalization) AdaLayerNormZeroSingle (class in tensorrt_llm.layers.normalization) adapter_id (tensorrt_llm.llmapi.LoRARequest property) add() (in module tensorrt_llm.functional) add_input() (tensorrt_llm.functional.Conditional method) add_note() (tensorrt_llm.llmapi.RequestError method) add_output() (tensorrt_llm.functional.Conditional method) add_sequence() (tensorrt_llm.runtime.KVCacheManager method) add_special_tokens (tensorrt_llm.llmapi.SamplingParams attribute) additional_context_outputs (tensorrt_llm.llmapi.CompletionOutput attribute) additional_generation_outputs (tensorrt_llm.llmapi.CompletionOutput attribute) additional_model_outputs (tensorrt_llm.llmapi.SamplingParams attribute) agent_hierarchy (tensorrt_llm.llmapi.SchedulingParams attribute) algorithm (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) (tensorrt_llm.llmapi.RocketSparseAttentionConfig attribute) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig attribute) alibi (tensorrt_llm.functional.PositionEmbeddingType attribute) alibi_with_scale (tensorrt_llm.functional.PositionEmbeddingType attribute) allgather() (in module tensorrt_llm.functional) allow_advanced_sampling (tensorrt_llm.llmapi.AutoDecodingConfig attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) allreduce() (in module tensorrt_llm.functional) allreduce_strategy (tensorrt_llm.llmapi.TorchLlmArgs attribute) AllReduceFusionOp (class in tensorrt_llm.functional) AllReduceParams (class in tensorrt_llm.functional) AllReduceStrategy (class in tensorrt_llm.functional) apply_batched_logits_processor (tensorrt_llm.llmapi.SamplingParams attribute) apply_llama3_scaling() (tensorrt_llm.functional.RopeEmbeddingUtils static method) apply_rotary_pos_emb() (tensorrt_llm.functional.RopeEmbeddingUtils static method) apply_rotary_pos_emb_chatglm() (tensorrt_llm.functional.RopeEmbeddingUtils static method) apply_rotary_pos_emb_cogvlm() (tensorrt_llm.functional.RopeEmbeddingUtils static method) arange() (in module tensorrt_llm.functional) aresult() (tensorrt_llm.llmapi.RequestOutput method) argmax() (in module tensorrt_llm.functional) args (tensorrt_llm.llmapi.RequestError attribute) as_integer_ratio() (tensorrt_llm.llmapi.DisaggScheduleStyle method) assert_valid_quant_algo() (tensorrt_llm.models.GemmaForCausalLM class method) assertion() (in module tensorrt_llm.functional) AsyncLLM (class in tensorrt_llm.llmapi) Attention (class in tensorrt_llm.layers.attention) attention_dp_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) attention_dp_events_gather_period_ms (tensorrt_llm.llmapi.KvCacheConfig attribute) attention_dp_rank (tensorrt_llm.llmapi.SchedulingParams attribute) attention_dp_relax (tensorrt_llm.llmapi.SchedulingParams attribute) AttentionDpConfig (class in tensorrt_llm.llmapi) AttentionDpConfig.Config (class in tensorrt_llm.llmapi) AttentionMaskParams (class in tensorrt_llm.layers.attention) AttentionMaskType (class in tensorrt_llm.functional) AttentionParams (class in tensorrt_llm.layers.attention) attn_backend (tensorrt_llm.llmapi.TorchLlmArgs attribute) attn_processors (tensorrt_llm.models.SD3Transformer2DModel property) audio_engine_dir (tensorrt_llm.runtime.MultimodalModelRunner property) AUTO (tensorrt_llm.functional.AllReduceStrategy attribute) (tensorrt_llm.models.SpeculativeDecodingMode attribute) AutoDecodingConfig (class in tensorrt_llm.llmapi) AutoDecodingConfig.Config (class in tensorrt_llm.llmapi) avg_decoded_tokens_per_iter (tensorrt_llm.llmapi.RequestOutput attribute) avg_pool2d() (in module tensorrt_llm.functional) AvgPool2d (class in tensorrt_llm.layers.pooling) axes (tensorrt_llm.functional.SliceInputType attribute) B backend (tensorrt_llm.llmapi.CacheTransceiverConfig attribute) (tensorrt_llm.llmapi.MoeConfig attribute) (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) bad (tensorrt_llm.llmapi.SamplingParams attribute) bad_token_ids (tensorrt_llm.llmapi.SamplingParams attribute) bad_words_list (tensorrt_llm.runtime.SamplingConfig attribute) BaichuanForCausalLM (class in tensorrt_llm.models) batch_size (tensorrt_llm.runtime.GenerationSession attribute) batch_sizes (tensorrt_llm.llmapi.CudaGraphConfig attribute) batch_wait_max_tokens_ratio (tensorrt_llm.llmapi.TorchLlmArgs attribute) batch_wait_timeout_iters (tensorrt_llm.llmapi.TorchLlmArgs attribute) batch_wait_timeout_ms (tensorrt_llm.llmapi.TorchLlmArgs attribute) batched_logits_processor (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) batching_type (tensorrt_llm.llmapi.TrtLlmArgs attribute) batching_wait_iters (tensorrt_llm.llmapi.AttentionDpConfig attribute) BatchingType (class in tensorrt_llm.llmapi) beam_search_diversity_rate (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) beam_width_array (tensorrt_llm.llmapi.SamplingParams attribute) begin_thinking_phase_token (tensorrt_llm.llmapi.MTPDecodingConfig attribute) bert_attention() (in module tensorrt_llm.functional) BertAttention (class in tensorrt_llm.layers.attention) BertForQuestionAnswering (class in tensorrt_llm.models) BertForSequenceClassification (class in tensorrt_llm.models) BertModel (class in tensorrt_llm.models) best_of (tensorrt_llm.llmapi.SamplingParams attribute) bidirectional (tensorrt_llm.functional.AttentionMaskType attribute) bidirectionalglm (tensorrt_llm.functional.AttentionMaskType attribute) bit_count() (tensorrt_llm.llmapi.DisaggScheduleStyle method) bit_length() (tensorrt_llm.llmapi.DisaggScheduleStyle method) blocksparse (tensorrt_llm.functional.AttentionMaskType attribute) BlockSparseAttnParams (class in tensorrt_llm.layers.attention) BloomForCausalLM (class in tensorrt_llm.models) BloomModel (class in tensorrt_llm.models) broadcast_helper() (in module tensorrt_llm.functional) buffer_allocated (tensorrt_llm.runtime.GenerationSession attribute) build_config (tensorrt_llm.llmapi.TrtLlmArgs attribute) BuildCacheConfig (class in tensorrt_llm.llmapi) BuildCacheConfig.Config (class in tensorrt_llm.llmapi) BuildConfig (class in tensorrt_llm.llmapi) BuildConfig.Config (class in tensorrt_llm.llmapi) C cache_root (tensorrt_llm.llmapi.BuildCacheConfig attribute) cache_transceiver_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) CacheTransceiverConfig (class in tensorrt_llm.llmapi) CacheTransceiverConfig.Config (class in tensorrt_llm.llmapi) calculate_speculative_resource() (tensorrt_llm.llmapi.LookaheadDecodingConfig method) calib_batch_size (tensorrt_llm.llmapi.CalibConfig attribute) calib_batches (tensorrt_llm.llmapi.CalibConfig attribute) calib_config (tensorrt_llm.llmapi.TrtLlmArgs attribute) calib_dataset (tensorrt_llm.llmapi.CalibConfig attribute) calib_max_seq_length (tensorrt_llm.llmapi.CalibConfig attribute) CalibConfig (class in tensorrt_llm.llmapi) CalibConfig.Config (class in tensorrt_llm.llmapi) candidate_metrics (tensorrt_llm.llmapi.RequestOutput attribute) capacity_scheduler_policy (tensorrt_llm.llmapi.SchedulerConfig attribute) CapacitySchedulerPolicy (class in tensorrt_llm.llmapi) capitalize() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) capture_num_tokens (tensorrt_llm.llmapi.TorchCompileConfig attribute) casefold() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) Cast (class in tensorrt_llm.layers.cast) cast() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) categorical_sample() (in module tensorrt_llm.functional) causal (tensorrt_llm.functional.AttentionMaskType attribute) center() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) chatglm (tensorrt_llm.functional.PositionEmbeddingType attribute) ChatGLMConfig (class in tensorrt_llm.models) ChatGLMForCausalLM (class in tensorrt_llm.models) ChatGLMGenerationSession (class in tensorrt_llm.runtime) ChatGLMModel (class in tensorrt_llm.models) check_config() (tensorrt_llm.models.DecoderModel method) (tensorrt_llm.models.DiT method) (tensorrt_llm.models.EncoderModel method) (tensorrt_llm.models.FalconForCausalLM method) (tensorrt_llm.models.MPTForCausalLM method) (tensorrt_llm.models.OPTForCausalLM method) (tensorrt_llm.models.PhiForCausalLM method) (tensorrt_llm.models.PretrainedModel method) check_eagle_choices() (tensorrt_llm.llmapi.Eagle3DecodingConfig method) (tensorrt_llm.llmapi.EagleDecodingConfig method) checkpoint_format (tensorrt_llm.llmapi.TorchLlmArgs attribute) checkpoint_loader (tensorrt_llm.llmapi.TorchLlmArgs attribute) choices() (tensorrt_llm.functional.PositionEmbeddingType static method) chunk() (in module tensorrt_llm.functional) ckpt_source (tensorrt_llm.llmapi.LoRARequest property) clamp_val (tensorrt_llm.llmapi.QuantConfig attribute) clear_logprob_params() (tensorrt_llm.llmapi.RequestOutput method) client_id (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Output attribute) clip() (in module tensorrt_llm.functional) CLIPVisionTransformer (class in tensorrt_llm.models) coerce_env_overrides_to_str() (tensorrt_llm.llmapi.TorchLlmArgs class method) (tensorrt_llm.llmapi.TrtLlmArgs class method) CogVLMAttention (class in tensorrt_llm.layers.attention) CogVLMConfig (class in tensorrt_llm.models) CogVLMForCausalLM (class in tensorrt_llm.models) CohereForCausalLM (class in tensorrt_llm.models) collect_and_bias() (tensorrt_llm.layers.linear.Linear method) (tensorrt_llm.layers.linear.LinearBase method) (tensorrt_llm.layers.linear.RowLinear method) collective_rpc() (tensorrt_llm.llmapi.AsyncLLM method) ColumnLinear (in module tensorrt_llm.layers.linear) CombinedTimestepLabelEmbeddings (class in tensorrt_llm.layers.embedding) CombinedTimestepTextProjEmbeddings (class in tensorrt_llm.layers.embedding) CompletionOutput (class in tensorrt_llm.llmapi) compute_relative_bias() (in module tensorrt_llm.layers.attention) concat() (in module tensorrt_llm.functional) Conditional (class in tensorrt_llm.functional) config_class (tensorrt_llm.models.BaichuanForCausalLM attribute) (tensorrt_llm.models.ChatGLMForCausalLM attribute) (tensorrt_llm.models.CogVLMForCausalLM attribute) (tensorrt_llm.models.CohereForCausalLM attribute) (tensorrt_llm.models.DbrxForCausalLM attribute) (tensorrt_llm.models.DeepseekForCausalLM attribute) (tensorrt_llm.models.DeepseekV2ForCausalLM attribute) (tensorrt_llm.models.EagleForCausalLM attribute) (tensorrt_llm.models.FalconForCausalLM attribute) (tensorrt_llm.models.GemmaForCausalLM attribute) (tensorrt_llm.models.GPTForCausalLM attribute) (tensorrt_llm.models.GPTJForCausalLM attribute) (tensorrt_llm.models.LLaMAForCausalLM attribute) (tensorrt_llm.models.MambaForCausalLM attribute) (tensorrt_llm.models.MedusaForCausalLm attribute) (tensorrt_llm.models.MLLaMAForCausalLM attribute) (tensorrt_llm.models.Phi3ForCausalLM attribute) (tensorrt_llm.models.PhiForCausalLM attribute) (tensorrt_llm.models.SD3Transformer2DModel attribute) conjugate() (tensorrt_llm.llmapi.DisaggScheduleStyle method) constant() (in module tensorrt_llm.functional) constant_to_tensor_() (in module tensorrt_llm.functional) constants_to_tensors_() (in module tensorrt_llm.functional) construct() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) context (tensorrt_llm.runtime.Session property) context_chunking_policy (tensorrt_llm.llmapi.SchedulerConfig attribute) CONTEXT_FIRST (tensorrt_llm.llmapi.DisaggScheduleStyle attribute) context_logits (tensorrt_llm.llmapi.RequestOutput attribute) (tensorrt_llm.llmapi.RequestOutput property) context_mem_size (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.Session property) context_parallel_size (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) ContextChunkingPolicy (class in tensorrt_llm.llmapi) Conv1d (class in tensorrt_llm.layers.conv) conv1d() (in module tensorrt_llm.functional) Conv2d (class in tensorrt_llm.layers.conv) conv2d() (in module tensorrt_llm.functional) Conv3d (class in tensorrt_llm.layers.conv) conv3d() (in module tensorrt_llm.functional) conv_kernel (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) conv_transpose2d() (in module tensorrt_llm.functional) convert_enable_disable() (tensorrt_llm.plugin.PluginConfig class method) convert_load_format() (tensorrt_llm.llmapi.TorchLlmArgs class method) ConvTranspose2d (class in tensorrt_llm.layers.conv) copy() (tensorrt_llm.llmapi.AttentionDpConfig method) (tensorrt_llm.llmapi.AutoDecodingConfig method) (tensorrt_llm.llmapi.BuildCacheConfig method) (tensorrt_llm.llmapi.BuildConfig method) (tensorrt_llm.llmapi.CacheTransceiverConfig method) (tensorrt_llm.llmapi.CalibConfig method) (tensorrt_llm.llmapi.CudaGraphConfig method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig method) (tensorrt_llm.llmapi.DFlashDecodingConfig method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig method) (tensorrt_llm.llmapi.DynamicBatchConfig method) (tensorrt_llm.llmapi.Eagle3DecodingConfig method) (tensorrt_llm.llmapi.EagleDecodingConfig method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig method) (tensorrt_llm.llmapi.KvCacheConfig method) (tensorrt_llm.llmapi.LookaheadDecodingConfig method) (tensorrt_llm.llmapi.MedusaDecodingConfig method) (tensorrt_llm.llmapi.MoeConfig method) (tensorrt_llm.llmapi.MTPDecodingConfig method) (tensorrt_llm.llmapi.NGramDecodingConfig method) (tensorrt_llm.llmapi.PARDDecodingConfig method) (tensorrt_llm.llmapi.PrometheusMetricsConfig method) (tensorrt_llm.llmapi.QuantConfig method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig method) (tensorrt_llm.llmapi.SADecodingConfig method) (tensorrt_llm.llmapi.SAEnhancerConfig method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig method) (tensorrt_llm.llmapi.SchedulerConfig method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig method) (tensorrt_llm.llmapi.TorchCompileConfig method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig method) copy_on_partial_reuse (tensorrt_llm.llmapi.KvCacheConfig attribute) cos() (in module tensorrt_llm.functional) count() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Output method) cp_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) cp_split_plugin() (in module tensorrt_llm.functional) cpp_e2e (tensorrt_llm.runtime.MultimodalModelRunner property) cpp_llm_only (tensorrt_llm.runtime.MultimodalModelRunner property) create_allreduce_plugin() (in module tensorrt_llm.functional) create_attention_const_params() (tensorrt_llm.layers.attention.Attention static method) create_fake_weight() (tensorrt_llm.functional.RopeEmbeddingUtils static method) create_runtime_defaults() (tensorrt_llm.models.PretrainedConfig static method) create_sinusoidal_positions() (tensorrt_llm.functional.RopeEmbeddingUtils static method) create_sinusoidal_positions_for_attention_plugin() (tensorrt_llm.functional.RopeEmbeddingUtils static method) create_sinusoidal_positions_for_cogvlm_attention_plugin() (tensorrt_llm.functional.RopeEmbeddingUtils static method) create_sinusoidal_positions_long_rope() (tensorrt_llm.functional.RopeEmbeddingUtils static method) create_sinusoidal_positions_long_rope_for_attention_plugin() (tensorrt_llm.functional.RopeEmbeddingUtils method) create_sinusoidal_positions_yarn() (tensorrt_llm.functional.RopeEmbeddingUtils static method) cropped_pos_embed() (tensorrt_llm.layers.embedding.SD3PatchEmbed method) cross_attention (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) cross_kv_cache_fraction (tensorrt_llm.llmapi.KvCacheConfig attribute) ctx_dp_rank (tensorrt_llm.llmapi.DisaggregatedParams attribute) ctx_info_endpoint (tensorrt_llm.llmapi.DisaggregatedParams attribute) ctx_request_id (tensorrt_llm.llmapi.DisaggregatedParams attribute) cuda_graph_cache_size (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig attribute) cuda_graph_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) cuda_graph_mode (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig attribute) (tensorrt_llm.runtime.GenerationSession attribute) cuda_stream_guard() (tensorrt_llm.runtime.GenerationSession method) cuda_stream_sync() (in module tensorrt_llm.functional) CudaGraphConfig (class in tensorrt_llm.llmapi) CudaGraphConfig.Config (class in tensorrt_llm.llmapi) cumsum() (in module tensorrt_llm.functional) cumulative_logprob (tensorrt_llm.llmapi.CompletionOutput attribute) custom_mask (tensorrt_llm.functional.AttentionMaskType attribute) custom_tokenizer (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) D data (tensorrt_llm.functional.SliceInputType attribute) DbrxConfig (class in tensorrt_llm.models) DbrxForCausalLM (class in tensorrt_llm.models) debug_mode (tensorrt_llm.runtime.GenerationSession attribute) debug_tensors_to_save (tensorrt_llm.runtime.GenerationSession attribute) decode() (tensorrt_llm.runtime.GenerationSession method) decode_batch() (tensorrt_llm.runtime.GenerationSession method) decode_duration_ms (tensorrt_llm.llmapi.KvCacheRetentionConfig property) decode_regular() (tensorrt_llm.runtime.GenerationSession method) decode_retention_priority (tensorrt_llm.llmapi.KvCacheRetentionConfig property) decode_stream() (tensorrt_llm.runtime.GenerationSession method) decode_words_list() (in module tensorrt_llm.runtime) DecoderModel (class in tensorrt_llm.models) decoding_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) decoding_type (tensorrt_llm.llmapi.AutoDecodingConfig attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) DeepseekForCausalLM (class in tensorrt_llm.models) DeepSeekSparseAttentionConfig (class in tensorrt_llm.llmapi) DeepSeekSparseAttentionConfig.Config (class in tensorrt_llm.llmapi) DeepseekV2Attention (class in tensorrt_llm.layers.attention) DeepseekV2ForCausalLM (class in tensorrt_llm.models) default_plugin_config() (tensorrt_llm.models.CogVLMForCausalLM method) (tensorrt_llm.models.LLaMAForCausalLM method) default_record_creator() (tensorrt_llm.llmapi.RequestOutput.PostprocWorker static method) deferred (tensorrt_llm.functional.PositionEmbeddingType attribute) denominator (tensorrt_llm.llmapi.DisaggScheduleStyle attribute) detokenize (tensorrt_llm.llmapi.SamplingParams attribute) device (tensorrt_llm.llmapi.CalibConfig attribute) (tensorrt_llm.runtime.GenerationSession attribute) DFlashDecodingConfig (class in tensorrt_llm.llmapi) DFlashDecodingConfig.Config (class in tensorrt_llm.llmapi) dict() (tensorrt_llm.llmapi.AttentionDpConfig method) (tensorrt_llm.llmapi.AutoDecodingConfig method) (tensorrt_llm.llmapi.BuildCacheConfig method) (tensorrt_llm.llmapi.BuildConfig method) (tensorrt_llm.llmapi.CacheTransceiverConfig method) (tensorrt_llm.llmapi.CalibConfig method) (tensorrt_llm.llmapi.CudaGraphConfig method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig method) (tensorrt_llm.llmapi.DFlashDecodingConfig method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig method) (tensorrt_llm.llmapi.DynamicBatchConfig method) (tensorrt_llm.llmapi.Eagle3DecodingConfig method) (tensorrt_llm.llmapi.EagleDecodingConfig method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig method) (tensorrt_llm.llmapi.KvCacheConfig method) (tensorrt_llm.llmapi.LookaheadDecodingConfig method) (tensorrt_llm.llmapi.MedusaDecodingConfig method) (tensorrt_llm.llmapi.MoeConfig method) (tensorrt_llm.llmapi.MTPDecodingConfig method) (tensorrt_llm.llmapi.NGramDecodingConfig method) (tensorrt_llm.llmapi.PARDDecodingConfig method) (tensorrt_llm.llmapi.PrometheusMetricsConfig method) (tensorrt_llm.llmapi.QuantConfig method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig method) (tensorrt_llm.llmapi.SADecodingConfig method) (tensorrt_llm.llmapi.SAEnhancerConfig method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig method) (tensorrt_llm.llmapi.SchedulerConfig method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig method) (tensorrt_llm.llmapi.TorchCompileConfig method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig method) DiffusersAttention (class in tensorrt_llm.layers.attention) DimRange (class in tensorrt_llm.functional) directory (tensorrt_llm.llmapi.KvCacheRetentionConfig property) disable (tensorrt_llm.functional.SideStreamIDType attribute) disable_finalize_fusion (tensorrt_llm.llmapi.MoeConfig attribute) disable_flashinfer_sampling (tensorrt_llm.llmapi.TorchLlmArgs attribute) disable_forward_chunking() (tensorrt_llm.models.SD3Transformer2DModel method) disable_overlap_scheduler (tensorrt_llm.llmapi.TorchLlmArgs attribute) disagg_request_id (tensorrt_llm.llmapi.DisaggregatedParams attribute) disaggregated_params (tensorrt_llm.llmapi.AsyncLLM property) (tensorrt_llm.llmapi.CompletionOutput attribute) (tensorrt_llm.llmapi.LLM attribute) (tensorrt_llm.llmapi.LLM property) (tensorrt_llm.llmapi.MultimodalEncoder property) (tensorrt_llm.llmapi.RequestOutput attribute) (tensorrt_llm.llmapi.RequestOutput property) (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Input attribute) (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Output attribute) DisaggregatedParams (class in tensorrt_llm.llmapi) DisaggScheduleStyle (class in tensorrt_llm.llmapi) DiT (class in tensorrt_llm.models) div() (in module tensorrt_llm.functional) do_tracing() (tensorrt_llm.llmapi.RequestOutput method) dora_plugin() (in module tensorrt_llm.functional) draft_len_schedule (tensorrt_llm.llmapi.AutoDecodingConfig attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) draft_tokens (tensorrt_llm.llmapi.DisaggregatedParams attribute) DRAFT_TOKENS_EXTERNAL (tensorrt_llm.models.SpeculativeDecodingMode attribute) drafter (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) DraftTargetDecodingConfig (class in tensorrt_llm.llmapi) DraftTargetDecodingConfig.Config (class in tensorrt_llm.llmapi) dry_run (tensorrt_llm.llmapi.BuildConfig attribute) dtype (tensorrt_llm.functional.Tensor property) (tensorrt_llm.llmapi.KvCacheConfig attribute) (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) (tensorrt_llm.runtime.ModelRunner property) (tensorrt_llm.runtime.ModelRunnerCpp property) (tensorrt_llm.runtime.TensorInfo attribute) dump_debug_buffers() (tensorrt_llm.runtime.GenerationSession method) duration_ms (tensorrt_llm.llmapi.KvCacheRetentionConfig.TokenRangeRetentionConfig property) dwdp_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) dynamic (tensorrt_llm.functional.RotaryScalingType attribute) dynamic_batch_config (tensorrt_llm.llmapi.SchedulerConfig attribute) dynamic_batch_moving_average_window (tensorrt_llm.llmapi.DynamicBatchConfig attribute) dynamic_tree_max_topK (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) DynamicBatchConfig (class in tensorrt_llm.llmapi) DynamicBatchConfig.Config (class in tensorrt_llm.llmapi) E e2e_request_latency_buckets (tensorrt_llm.llmapi.PrometheusMetricsConfig attribute) EAGLE (tensorrt_llm.models.SpeculativeDecodingMode attribute) eagle3_layers_to_capture (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) eagle3_model_arch (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) eagle3_one_model (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) Eagle3DecodingConfig (class in tensorrt_llm.llmapi) Eagle3DecodingConfig.Config (class in tensorrt_llm.llmapi) eagle_choices (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) EagleDecodingConfig (class in tensorrt_llm.llmapi) EagleDecodingConfig.Config (class in tensorrt_llm.llmapi) EagleForCausalLM (class in tensorrt_llm.models) early_stop_criteria() (tensorrt_llm.runtime.GenerationSession method) early_stopping (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) einsum() (in module tensorrt_llm.functional) elementwise_binary() (in module tensorrt_llm.functional) Embedding (class in tensorrt_llm.layers.embedding) embedding() (in module tensorrt_llm.functional) embedding_bias (tensorrt_llm.llmapi.SamplingParams attribute) embedding_parallel_mode (tensorrt_llm.llmapi.TrtLlmArgs attribute) enable_attention_dp (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) enable_autotuner (tensorrt_llm.llmapi.TorchLlmArgs attribute) enable_balance (tensorrt_llm.llmapi.AttentionDpConfig attribute) enable_batch_size_tuning (tensorrt_llm.llmapi.DynamicBatchConfig attribute) enable_block_reuse (tensorrt_llm.llmapi.KvCacheConfig attribute) enable_build_cache (tensorrt_llm.llmapi.TrtLlmArgs attribute) enable_chunked_prefill (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) enable_context_fmha_fp32_acc (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig attribute) enable_debug_output (tensorrt_llm.llmapi.BuildConfig attribute) enable_early_first_token_response (tensorrt_llm.llmapi.TorchLlmArgs attribute) enable_energy_metrics (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) enable_forward_chunking() (tensorrt_llm.models.SD3Transformer2DModel method) enable_fullgraph (tensorrt_llm.llmapi.TorchCompileConfig attribute) enable_global_pool (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SAEnhancerConfig attribute) enable_heuristic_topk (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) enable_inductor (tensorrt_llm.llmapi.TorchCompileConfig attribute) enable_iter_perf_stats (tensorrt_llm.llmapi.TorchLlmArgs attribute) enable_iter_req_stats (tensorrt_llm.llmapi.TorchLlmArgs attribute) enable_kv_cache_aware_routing (tensorrt_llm.llmapi.AttentionDpConfig attribute) enable_layerwise_nvtx_marker (tensorrt_llm.llmapi.TorchLlmArgs attribute) enable_lm_head_tp_in_adp (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) enable_lora (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) enable_max_num_tokens_tuning (tensorrt_llm.llmapi.DynamicBatchConfig attribute) enable_min_latency (tensorrt_llm.llmapi.TorchLlmArgs attribute) enable_padding (tensorrt_llm.llmapi.CudaGraphConfig attribute) enable_partial_reuse (tensorrt_llm.llmapi.KvCacheConfig attribute) enable_piecewise_cuda_graph (tensorrt_llm.llmapi.TorchCompileConfig attribute) enable_prompt_adapter (tensorrt_llm.llmapi.TrtLlmArgs attribute) enable_resource_governor (tensorrt_llm.llmapi.TorchLlmArgs attribute) enable_speculative_beam_history_d2h (tensorrt_llm.llmapi.TorchLlmArgs attribute) enable_tqdm (tensorrt_llm.llmapi.TrtLlmArgs attribute) enable_userbuffers (tensorrt_llm.llmapi.TorchCompileConfig attribute) EncDecModelRunner (class in tensorrt_llm.runtime) encode() (tensorrt_llm.llmapi.AsyncLLM method) (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.LLM method) (tensorrt_llm.llmapi.MultimodalEncoder method) (tensorrt_llm.llmapi.QuantAlgo method) encode_only (tensorrt_llm.llmapi.TorchLlmArgs attribute) encoder_run() (tensorrt_llm.runtime.EncDecModelRunner method) EncoderModel (class in tensorrt_llm.models) end_id (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) end_thinking_phase_token (tensorrt_llm.llmapi.MTPDecodingConfig attribute) endswith() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) engine (tensorrt_llm.runtime.Session property) engine_inspector (tensorrt_llm.runtime.GenerationSession property) env_overrides (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) eq() (in module tensorrt_llm.functional) EQUAL_PROGRESS (tensorrt_llm.llmapi.ContextChunkingPolicy attribute) error (tensorrt_llm.llmapi.RequestOutput attribute) (tensorrt_llm.llmapi.RequestOutput property) event_buffer_max_size (tensorrt_llm.llmapi.KvCacheConfig attribute) exclude_input_from_output (tensorrt_llm.llmapi.SamplingParams attribute) exclude_modules (tensorrt_llm.llmapi.QuantConfig attribute) exp() (in module tensorrt_llm.functional) expand() (in module tensorrt_llm.functional) expand_dims() (in module tensorrt_llm.functional) expand_dims_like() (in module tensorrt_llm.functional) expand_mask() (in module tensorrt_llm.functional) expandtabs() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) EXPLICIT_DRAFT_TOKENS (tensorrt_llm.models.SpeculativeDecodingMode attribute) extended_runtime_perf_knob_config (tensorrt_llm.llmapi.TrtLlmArgs attribute) ExtendedRuntimePerfKnobConfig (class in tensorrt_llm.llmapi) ExtendedRuntimePerfKnobConfig.Config (class in tensorrt_llm.llmapi) extra (tensorrt_llm.llmapi.AttentionDpConfig.Config attribute) (tensorrt_llm.llmapi.AutoDecodingConfig.Config attribute) (tensorrt_llm.llmapi.BuildCacheConfig.Config attribute) (tensorrt_llm.llmapi.BuildConfig.Config attribute) (tensorrt_llm.llmapi.CacheTransceiverConfig.Config attribute) (tensorrt_llm.llmapi.CalibConfig.Config attribute) (tensorrt_llm.llmapi.CudaGraphConfig.Config attribute) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig.Config attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig.Config attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig.Config attribute) (tensorrt_llm.llmapi.DynamicBatchConfig.Config attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig.Config attribute) (tensorrt_llm.llmapi.EagleDecodingConfig.Config attribute) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig.Config attribute) (tensorrt_llm.llmapi.KvCacheConfig.Config attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig.Config attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig.Config attribute) (tensorrt_llm.llmapi.MoeConfig.Config attribute) (tensorrt_llm.llmapi.MTPDecodingConfig.Config attribute) (tensorrt_llm.llmapi.NGramDecodingConfig.Config attribute) (tensorrt_llm.llmapi.PARDDecodingConfig.Config attribute) (tensorrt_llm.llmapi.PrometheusMetricsConfig.Config attribute) (tensorrt_llm.llmapi.QuantConfig.Config attribute) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig.Config attribute) (tensorrt_llm.llmapi.RocketSparseAttentionConfig.Config attribute) (tensorrt_llm.llmapi.SADecodingConfig.Config attribute) (tensorrt_llm.llmapi.SAEnhancerConfig.Config attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig.Config attribute) (tensorrt_llm.llmapi.SchedulerConfig.Config attribute) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig.Config attribute) (tensorrt_llm.llmapi.TorchCompileConfig.Config attribute) (tensorrt_llm.llmapi.TorchLlmArgs.Config attribute) (tensorrt_llm.llmapi.TrtLlmArgs.Config attribute) (tensorrt_llm.llmapi.UserProvidedDecodingConfig.Config attribute) extra_resource_managers (tensorrt_llm.llmapi.TorchLlmArgs property) F fail_fast_on_attention_window_too_large (tensorrt_llm.llmapi.TrtLlmArgs attribute) FalconConfig (class in tensorrt_llm.models) FalconForCausalLM (class in tensorrt_llm.models) FalconModel (class in tensorrt_llm.models) fast_build (tensorrt_llm.llmapi.TrtLlmArgs attribute) fc_gate() (tensorrt_llm.layers.mlp.FusedGatedMLP method) fc_gate_dora() (in module tensorrt_llm.layers.mlp) fc_gate_lora() (in module tensorrt_llm.layers.mlp) fc_gate_plugin() (tensorrt_llm.layers.mlp.FusedGatedMLP method) field_name (tensorrt_llm.llmapi.TorchLlmArgs attribute), [1] (tensorrt_llm.llmapi.TrtLlmArgs attribute) file_prefix (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) fill_attention_const_params_for_long_rope() (tensorrt_llm.layers.attention.AttentionParams method) fill_attention_const_params_for_rope() (tensorrt_llm.layers.attention.AttentionParams method) fill_attention_params() (tensorrt_llm.layers.attention.Attention static method) fill_none_tensor_list() (tensorrt_llm.layers.attention.KeyValueCacheParams method) fill_value (tensorrt_llm.functional.SliceInputType attribute) filter_medusa_logits() (tensorrt_llm.runtime.GenerationSession method) finalize_decoder() (tensorrt_llm.runtime.GenerationSession method) find() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) find_best_medusa_path() (tensorrt_llm.runtime.GenerationSession method) finish_reason (tensorrt_llm.llmapi.CompletionOutput attribute) finished (tensorrt_llm.llmapi.RequestOutput attribute) (tensorrt_llm.llmapi.RequestOutput property) FIRST_COME_FIRST_SERVED (tensorrt_llm.llmapi.ContextChunkingPolicy attribute) first_gen_log_probs (tensorrt_llm.llmapi.DisaggregatedParams attribute) first_gen_logits (tensorrt_llm.llmapi.DisaggregatedParams attribute) first_gen_tokens (tensorrt_llm.llmapi.DisaggregatedParams attribute) first_layer (tensorrt_llm.runtime.GenerationSession property) flatten() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) flip() (in module tensorrt_llm.functional) floordiv() (in module tensorrt_llm.functional) for_each_rank() (tensorrt_llm.models.PretrainedConfig method) FORCE_CHUNK (tensorrt_llm.llmapi.ContextChunkingPolicy attribute) force_dynamic_quantization (tensorrt_llm.llmapi.TorchLlmArgs attribute) force_num_profiles (tensorrt_llm.llmapi.BuildConfig attribute) format() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) format_map() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) forward() (tensorrt_llm.layers.activation.Mish method) (tensorrt_llm.layers.attention.Attention method) (tensorrt_llm.layers.attention.BertAttention method) (tensorrt_llm.layers.attention.CogVLMAttention method) (tensorrt_llm.layers.attention.DeepseekV2Attention method) (tensorrt_llm.layers.attention.DiffusersAttention method) (tensorrt_llm.layers.cast.Cast method) (tensorrt_llm.layers.conv.Conv1d method) (tensorrt_llm.layers.conv.Conv2d method) (tensorrt_llm.layers.conv.Conv3d method) (tensorrt_llm.layers.conv.ConvTranspose2d method) (tensorrt_llm.layers.embedding.CombinedTimestepLabelEmbeddings method) (tensorrt_llm.layers.embedding.CombinedTimestepTextProjEmbeddings method) (tensorrt_llm.layers.embedding.Embedding method) (tensorrt_llm.layers.embedding.LabelEmbedding method) (tensorrt_llm.layers.embedding.PixArtAlphaTextProjection method) (tensorrt_llm.layers.embedding.PromptTuningEmbedding method) (tensorrt_llm.layers.embedding.SD3PatchEmbed method) (tensorrt_llm.layers.embedding.TimestepEmbedding method) (tensorrt_llm.layers.embedding.Timesteps method) (tensorrt_llm.layers.linear.LinearBase method) (tensorrt_llm.layers.mlp.FusedGatedMLP method) (tensorrt_llm.layers.mlp.GatedMLP method) (tensorrt_llm.layers.mlp.LinearActivation method) (tensorrt_llm.layers.mlp.LinearApproximateGELU method) (tensorrt_llm.layers.mlp.LinearGEGLU method) (tensorrt_llm.layers.mlp.LinearGELU method) (tensorrt_llm.layers.mlp.LinearSwiGLU method) (tensorrt_llm.layers.mlp.MLP method) (tensorrt_llm.layers.normalization.AdaLayerNorm method) (tensorrt_llm.layers.normalization.AdaLayerNormContinuous method) (tensorrt_llm.layers.normalization.AdaLayerNormZero method) (tensorrt_llm.layers.normalization.AdaLayerNormZeroSingle method) (tensorrt_llm.layers.normalization.GroupNorm method) (tensorrt_llm.layers.normalization.LayerNorm method) (tensorrt_llm.layers.normalization.RmsNorm method) (tensorrt_llm.layers.normalization.SD35AdaLayerNormZeroX method) (tensorrt_llm.layers.pooling.AvgPool2d method) (tensorrt_llm.models.BertForQuestionAnswering method) (tensorrt_llm.models.BertForSequenceClassification method) (tensorrt_llm.models.BertModel method) (tensorrt_llm.models.BloomModel method) (tensorrt_llm.models.ChatGLMModel method) (tensorrt_llm.models.CLIPVisionTransformer method) (tensorrt_llm.models.DecoderModel method) (tensorrt_llm.models.DiT method) (tensorrt_llm.models.EagleForCausalLM method) (tensorrt_llm.models.EncoderModel method) (tensorrt_llm.models.FalconModel method) (tensorrt_llm.models.GPTJModel method) (tensorrt_llm.models.GPTModel method) (tensorrt_llm.models.GPTNeoXModel method) (tensorrt_llm.models.LLaMAModel method) (tensorrt_llm.models.LlavaNextVisionWrapper method) (tensorrt_llm.models.MambaForCausalLM method) (tensorrt_llm.models.MLLaMAForCausalLM method) (tensorrt_llm.models.MPTModel method) (tensorrt_llm.models.OPTModel method) (tensorrt_llm.models.Phi3Model method) (tensorrt_llm.models.PhiModel method) (tensorrt_llm.models.RecurrentGemmaForCausalLM method) (tensorrt_llm.models.SD3Transformer2DModel method) (tensorrt_llm.models.WhisperEncoder method) forward_with_cfg() (tensorrt_llm.models.DiT method) forward_without_cfg() (tensorrt_llm.models.DiT method) FP8 (tensorrt_llm.llmapi.QuantAlgo attribute) FP8_BLOCK_SCALES (tensorrt_llm.llmapi.QuantAlgo attribute) FP8_PER_CHANNEL_PER_TOKEN (tensorrt_llm.llmapi.QuantAlgo attribute) free_gpu_memory_fraction (tensorrt_llm.llmapi.KvCacheConfig attribute) frequency_penalty (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) from_arguments() (tensorrt_llm.models.SpeculativeDecodingMode static method) (tensorrt_llm.plugin.PluginConfig class method) from_bytes() (tensorrt_llm.llmapi.DisaggScheduleStyle class method) from_checkpoint() (tensorrt_llm.models.PretrainedConfig class method) (tensorrt_llm.models.PretrainedModel class method) from_config() (tensorrt_llm.models.PretrainedModel class method) from_dict() (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.models.PretrainedConfig class method) from_dir() (tensorrt_llm.runtime.ModelRunner class method) (tensorrt_llm.runtime.ModelRunnerCpp class method) from_engine() (tensorrt_llm.runtime.EncDecModelRunner class method) (tensorrt_llm.runtime.ModelRunner class method) (tensorrt_llm.runtime.Session static method) from_hugging_face() (tensorrt_llm.models.BaichuanForCausalLM class method) (tensorrt_llm.models.ChatGLMConfig class method) (tensorrt_llm.models.ChatGLMForCausalLM class method) (tensorrt_llm.models.CogVLMForCausalLM class method) (tensorrt_llm.models.CohereForCausalLM class method) (tensorrt_llm.models.DeepseekForCausalLM class method) (tensorrt_llm.models.DeepseekV2ForCausalLM class method) (tensorrt_llm.models.EagleForCausalLM class method) (tensorrt_llm.models.FalconConfig class method) (tensorrt_llm.models.FalconForCausalLM class method) (tensorrt_llm.models.GemmaConfig class method) (tensorrt_llm.models.GemmaForCausalLM class method) (tensorrt_llm.models.GPTConfig class method) (tensorrt_llm.models.GPTForCausalLM class method) (tensorrt_llm.models.GPTJConfig class method) (tensorrt_llm.models.GPTJForCausalLM class method) (tensorrt_llm.models.LLaMAConfig class method) (tensorrt_llm.models.LLaMAForCausalLM class method) (tensorrt_llm.models.LlavaNextVisionConfig class method) (tensorrt_llm.models.LlavaNextVisionWrapper class method) (tensorrt_llm.models.MambaForCausalLM class method) (tensorrt_llm.models.MedusaConfig class method) (tensorrt_llm.models.MedusaForCausalLm class method) (tensorrt_llm.models.MLLaMAForCausalLM class method) (tensorrt_llm.models.Phi3ForCausalLM class method) (tensorrt_llm.models.PhiForCausalLM class method) from_json_file() (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.models.PretrainedConfig class method) from_meta_ckpt() (tensorrt_llm.models.LLaMAConfig class method) (tensorrt_llm.models.LLaMAForCausalLM class method) from_model_config_cpp() (tensorrt_llm.runtime.ModelConfig class method) from_nemo() (tensorrt_llm.models.GPTConfig class method) (tensorrt_llm.models.GPTForCausalLM class method) from_orm() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) from_pretrained() (tensorrt_llm.models.SD3Transformer2DModel class method) from_pybind() (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) from_serialized_engine() (tensorrt_llm.runtime.Session static method) from_string() (tensorrt_llm.functional.PositionEmbeddingType static method) (tensorrt_llm.functional.RotaryScalingType static method) from_yaml() (tensorrt_llm.llmapi.TorchLlmArgs class method) (tensorrt_llm.llmapi.TrtLlmArgs class method) fuse_qkv_projections() (tensorrt_llm.models.SD3Transformer2DModel method) FusedGatedMLP (class in tensorrt_llm.layers.mlp) (tensorrt_llm.functional.MLPType attribute) G garbage_collection_gen0_threshold (tensorrt_llm.llmapi.TorchLlmArgs attribute) GatedMLP (class in tensorrt_llm.layers.mlp) (tensorrt_llm.functional.MLPType attribute) gather() (in module tensorrt_llm.functional) gather_context_logits (tensorrt_llm.llmapi.BuildConfig attribute) (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) (tensorrt_llm.runtime.ModelRunner property) (tensorrt_llm.runtime.ModelRunnerCpp property) gather_generation_logits (tensorrt_llm.llmapi.BuildConfig attribute) (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) (tensorrt_llm.runtime.ModelRunner property) (tensorrt_llm.runtime.ModelRunnerCpp property) gather_last_token_logits() (in module tensorrt_llm.functional) gather_nd() (in module tensorrt_llm.functional) gegelu() (in module tensorrt_llm.functional) geglu() (in module tensorrt_llm.functional) gelu() (in module tensorrt_llm.functional) gemm_allreduce() (in module tensorrt_llm.functional) gemm_allreduce_plugin (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) gemm_swiglu() (in module tensorrt_llm.functional) GEMMA2_ADDED_FIELDS (tensorrt_llm.models.GemmaConfig attribute) gemma2_config() (tensorrt_llm.models.GemmaConfig method) GEMMA3_ADDED_FIELDS (tensorrt_llm.models.GemmaConfig attribute) gemma3_config() (tensorrt_llm.models.GemmaConfig method) GEMMA_ADDED_FIELDS (tensorrt_llm.models.GemmaConfig attribute) GemmaConfig (class in tensorrt_llm.models) GemmaForCausalLM (class in tensorrt_llm.models) generate() (tensorrt_llm.llmapi.AsyncLLM method) (tensorrt_llm.llmapi.LLM method) (tensorrt_llm.llmapi.MultimodalEncoder method) (tensorrt_llm.runtime.EncDecModelRunner method) (tensorrt_llm.runtime.ModelRunner method) (tensorrt_llm.runtime.ModelRunnerCpp method) (tensorrt_llm.runtime.MultimodalModelRunner method) (tensorrt_llm.runtime.QWenForCausalLMGenerationSession method) generate_alibi_biases() (in module tensorrt_llm.functional) generate_alibi_slopes() (in module tensorrt_llm.functional) generate_async() (tensorrt_llm.llmapi.AsyncLLM method) (tensorrt_llm.llmapi.LLM method) (tensorrt_llm.llmapi.MultimodalEncoder method) generate_logn_scaling() (in module tensorrt_llm.functional) GENERATION_FIRST (tensorrt_llm.llmapi.DisaggScheduleStyle attribute) generation_logits (tensorrt_llm.llmapi.CompletionOutput attribute) GenerationSequence (class in tensorrt_llm.runtime) GenerationSession (class in tensorrt_llm.runtime) get_1d_sincos_pos_embed_from_grid() (in module tensorrt_llm.layers.embedding) get_2d_sincos_pos_embed() (in module tensorrt_llm.layers.embedding) get_2d_sincos_pos_embed_from_grid() (in module tensorrt_llm.layers.embedding) get_audio_features() (tensorrt_llm.runtime.MultimodalModelRunner method) get_batch_idx() (tensorrt_llm.runtime.GenerationSequence method) get_block_offsets() (tensorrt_llm.runtime.KVCacheManager method) get_comm() (tensorrt_llm.llmapi.MpiCommSession method) get_config_group() (tensorrt_llm.models.PretrainedConfig method) get_context_phase_params() (tensorrt_llm.llmapi.DisaggregatedParams method) get_executor_config() (tensorrt_llm.llmapi.TorchLlmArgs method) get_first_past_key_value() (tensorrt_llm.layers.attention.KeyValueCacheParams method) get_hf_config() (tensorrt_llm.models.GemmaConfig static method) get_indices_block_size() (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig method) get_kv_cache_events() (tensorrt_llm.llmapi.AsyncLLM method) (tensorrt_llm.llmapi.LLM method) (tensorrt_llm.llmapi.MultimodalEncoder method) get_kv_cache_events_async() (tensorrt_llm.llmapi.AsyncLLM method) (tensorrt_llm.llmapi.LLM method) (tensorrt_llm.llmapi.MultimodalEncoder method) get_next_medusa_tokens() (tensorrt_llm.runtime.GenerationSession method) get_num_heads_kv() (tensorrt_llm.runtime.GenerationSession method) get_parent() (tensorrt_llm.functional.Tensor method) get_pybind_enum_fields() (tensorrt_llm.llmapi.CacheTransceiverConfig static method) (tensorrt_llm.llmapi.DynamicBatchConfig static method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig static method) (tensorrt_llm.llmapi.KvCacheConfig static method) (tensorrt_llm.llmapi.LookaheadDecodingConfig static method) (tensorrt_llm.llmapi.SchedulerConfig static method) get_pybind_variable_fields() (tensorrt_llm.llmapi.CacheTransceiverConfig static method) (tensorrt_llm.llmapi.DynamicBatchConfig static method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig static method) (tensorrt_llm.llmapi.KvCacheConfig static method) (tensorrt_llm.llmapi.LookaheadDecodingConfig static method) (tensorrt_llm.llmapi.SchedulerConfig static method) get_request_type() (tensorrt_llm.llmapi.DisaggregatedParams method) get_rope_index() (tensorrt_llm.runtime.MultimodalModelRunner method) get_runtime_sizes() (tensorrt_llm.llmapi.TorchLlmArgs method) (tensorrt_llm.llmapi.TrtLlmArgs method) get_seq_idx() (tensorrt_llm.runtime.GenerationSequence method) get_stats() (tensorrt_llm.llmapi.AsyncLLM method) (tensorrt_llm.llmapi.LLM method) (tensorrt_llm.llmapi.MultimodalEncoder method) get_stats_async() (tensorrt_llm.llmapi.AsyncLLM method) (tensorrt_llm.llmapi.LLM method) (tensorrt_llm.llmapi.MultimodalEncoder method) get_timestep_embedding() (in module tensorrt_llm.layers.embedding) get_users() (tensorrt_llm.functional.Tensor method) get_visual_features() (tensorrt_llm.runtime.MultimodalModelRunner method) get_weight() (tensorrt_llm.layers.linear.LinearBase method) global_pool_size (tensorrt_llm.llmapi.SADecodingConfig attribute) gpt_attention() (in module tensorrt_llm.functional) gpt_attention_plugin (tensorrt_llm.runtime.ModelConfig attribute) GPTConfig (class in tensorrt_llm.models) GPTForCausalLM (class in tensorrt_llm.models) GPTJConfig (class in tensorrt_llm.models) GPTJForCausalLM (class in tensorrt_llm.models) GPTJModel (class in tensorrt_llm.models) GPTModel (class in tensorrt_llm.models) GPTNeoXForCausalLM (class in tensorrt_llm.models) GPTNeoXModel (class in tensorrt_llm.models) gpu_weights_percent (tensorrt_llm.runtime.ModelConfig attribute) gpus_per_node (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) grammar (tensorrt_llm.llmapi.GuidedDecodingParams attribute) greedy_sampling (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) group_norm() (in module tensorrt_llm.functional) group_size (tensorrt_llm.llmapi.QuantConfig attribute) GroupNorm (class in tensorrt_llm.layers.normalization) (tensorrt_llm.functional.LayerNormType attribute) gt() (in module tensorrt_llm.functional) GUARANTEED_NO_EVICT (tensorrt_llm.llmapi.CapacitySchedulerPolicy attribute) guided_decoding (tensorrt_llm.llmapi.SamplingParams attribute) guided_decoding_backend (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) GuidedDecodingParams (class in tensorrt_llm.llmapi) H handle_per_step() (tensorrt_llm.runtime.GenerationSession method) has_affine() (tensorrt_llm.functional.AllReduceParams method) has_bias() (tensorrt_llm.functional.AllReduceParams method) has_config_group() (tensorrt_llm.models.PretrainedConfig method) has_position_embedding (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) has_scale() (tensorrt_llm.functional.AllReduceParams method) has_token_type_embedding (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) has_zero_point (tensorrt_llm.llmapi.QuantConfig attribute) head_size (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) hidden_size (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) (tensorrt_llm.runtime.ModelRunner property) (tensorrt_llm.runtime.ModelRunnerCpp property) host_cache_size (tensorrt_llm.llmapi.KvCacheConfig attribute) I identity() (in module tensorrt_llm.functional) ignore_eos (tensorrt_llm.llmapi.SamplingParams attribute) imag (tensorrt_llm.llmapi.DisaggScheduleStyle attribute) include_stop_str_in_output (tensorrt_llm.llmapi.SamplingParams attribute) index (tensorrt_llm.llmapi.CompletionOutput attribute) index() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Output method) index_head_dim (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) index_n_heads (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) index_select() (in module tensorrt_llm.functional) index_topk (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) indexer_k_dtype (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) indexer_max_chunk_size (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) indexer_rope_interleave (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) infer_shapes() (tensorrt_llm.runtime.Session method) INFLIGHT (tensorrt_llm.llmapi.BatchingType attribute) init_audio_encoder() (tensorrt_llm.runtime.MultimodalModelRunner method) init_build_config() (tensorrt_llm.llmapi.TrtLlmArgs method) init_image_encoder() (tensorrt_llm.runtime.MultimodalModelRunner method) init_llm() (tensorrt_llm.runtime.MultimodalModelRunner method) init_processor() (tensorrt_llm.runtime.MultimodalModelRunner method) init_tokenizer() (tensorrt_llm.runtime.MultimodalModelRunner method) input_timing_cache (tensorrt_llm.llmapi.BuildConfig attribute) INT8 (tensorrt_llm.llmapi.QuantAlgo attribute) int_clip() (in module tensorrt_llm.functional) interpolate() (in module tensorrt_llm.functional) is_alibi() (tensorrt_llm.functional.PositionEmbeddingType method) is_comm_session() (tensorrt_llm.llmapi.MpiCommSession method) is_deferred() (tensorrt_llm.functional.PositionEmbeddingType method) is_dynamic() (tensorrt_llm.functional.Tensor method) is_final (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Output attribute) is_gated_activation() (in module tensorrt_llm.functional) is_gemma_2 (tensorrt_llm.models.GemmaConfig property) is_gemma_3 (tensorrt_llm.models.GemmaConfig property) is_integer() (tensorrt_llm.llmapi.DisaggScheduleStyle method) is_keep_all (tensorrt_llm.llmapi.NGramDecodingConfig attribute) is_linear_tree (tensorrt_llm.llmapi.AutoDecodingConfig property) (tensorrt_llm.llmapi.DFlashDecodingConfig property) (tensorrt_llm.llmapi.DraftTargetDecodingConfig property) (tensorrt_llm.llmapi.Eagle3DecodingConfig property) (tensorrt_llm.llmapi.EagleDecodingConfig property) (tensorrt_llm.llmapi.LookaheadDecodingConfig property) (tensorrt_llm.llmapi.MedusaDecodingConfig property) (tensorrt_llm.llmapi.MTPDecodingConfig property) (tensorrt_llm.llmapi.NGramDecodingConfig property) (tensorrt_llm.llmapi.PARDDecodingConfig property) (tensorrt_llm.llmapi.SADecodingConfig property) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig property) (tensorrt_llm.llmapi.UserProvidedDecodingConfig property) is_medusa_mode (tensorrt_llm.runtime.GenerationSession property) is_module_excluded_from_quantization() (tensorrt_llm.llmapi.QuantConfig method) is_mrope() (tensorrt_llm.functional.PositionEmbeddingType method) is_public_pool (tensorrt_llm.llmapi.NGramDecodingConfig attribute) is_redrafter_mode (tensorrt_llm.runtime.GenerationSession property) is_rope() (tensorrt_llm.functional.PositionEmbeddingType method) is_trt_wrapper() (tensorrt_llm.functional.Tensor method) is_use_oldest (tensorrt_llm.llmapi.NGramDecodingConfig attribute) is_valid() (tensorrt_llm.functional.MoEAllReduceParams method) (tensorrt_llm.layers.attention.AttentionParams method) (tensorrt_llm.layers.attention.KeyValueCacheParams method) is_valid_cross_attn() (tensorrt_llm.layers.attention.AttentionParams method) isalnum() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) isalpha() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) isascii() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) isdecimal() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) isdigit() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) isidentifier() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) islower() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) isnumeric() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) isprintable() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) isspace() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) istitle() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) isupper() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) iter_stats_max_iterations (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) iteration_stats_interval (tensorrt_llm.llmapi.KvCacheConfig attribute) J join() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) joint_attn_forward() (tensorrt_llm.layers.attention.DiffusersAttention method) json (tensorrt_llm.llmapi.GuidedDecodingParams attribute) json() (tensorrt_llm.llmapi.AttentionDpConfig method) (tensorrt_llm.llmapi.AutoDecodingConfig method) (tensorrt_llm.llmapi.BuildCacheConfig method) (tensorrt_llm.llmapi.BuildConfig method) (tensorrt_llm.llmapi.CacheTransceiverConfig method) (tensorrt_llm.llmapi.CalibConfig method) (tensorrt_llm.llmapi.CudaGraphConfig method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig method) (tensorrt_llm.llmapi.DFlashDecodingConfig method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig method) (tensorrt_llm.llmapi.DynamicBatchConfig method) (tensorrt_llm.llmapi.Eagle3DecodingConfig method) (tensorrt_llm.llmapi.EagleDecodingConfig method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig method) (tensorrt_llm.llmapi.KvCacheConfig method) (tensorrt_llm.llmapi.LookaheadDecodingConfig method) (tensorrt_llm.llmapi.MedusaDecodingConfig method) (tensorrt_llm.llmapi.MoeConfig method) (tensorrt_llm.llmapi.MTPDecodingConfig method) (tensorrt_llm.llmapi.NGramDecodingConfig method) (tensorrt_llm.llmapi.PARDDecodingConfig method) (tensorrt_llm.llmapi.PrometheusMetricsConfig method) (tensorrt_llm.llmapi.QuantConfig method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig method) (tensorrt_llm.llmapi.SADecodingConfig method) (tensorrt_llm.llmapi.SAEnhancerConfig method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig method) (tensorrt_llm.llmapi.SchedulerConfig method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig method) (tensorrt_llm.llmapi.TorchCompileConfig method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig method) json_object (tensorrt_llm.llmapi.GuidedDecodingParams attribute) K kernel_size (tensorrt_llm.llmapi.RocketSparseAttentionConfig attribute) KeyValueCacheParams (class in tensorrt_llm.layers.attention) kt_cache_dtype (tensorrt_llm.llmapi.RocketSparseAttentionConfig attribute) kv_cache_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) kv_cache_quant_algo (tensorrt_llm.llmapi.QuantConfig attribute) kv_cache_routing_fair_share_multiplier (tensorrt_llm.llmapi.AttentionDpConfig attribute) kv_cache_routing_load_balance_weight (tensorrt_llm.llmapi.AttentionDpConfig attribute) kv_cache_routing_match_rate_threshold (tensorrt_llm.llmapi.AttentionDpConfig attribute) kv_cache_type (tensorrt_llm.llmapi.BuildConfig attribute) (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) kv_connector_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) kv_dtype (tensorrt_llm.models.PretrainedConfig property) kv_transfer_sender_future_timeout_ms (tensorrt_llm.llmapi.CacheTransceiverConfig attribute) kv_transfer_timeout_ms (tensorrt_llm.llmapi.CacheTransceiverConfig attribute) KvCacheConfig (class in tensorrt_llm.llmapi) KvCacheConfig.Config (class in tensorrt_llm.llmapi) KVCacheManager (class in tensorrt_llm.runtime) KvCacheRetentionConfig (class in tensorrt_llm.llmapi) KvCacheRetentionConfig.TokenRangeRetentionConfig (class in tensorrt_llm.llmapi) L LabelEmbedding (class in tensorrt_llm.layers.embedding) language_adapter_config (tensorrt_llm.runtime.ModelConfig attribute) last_layer (tensorrt_llm.runtime.GenerationSession property) LAST_PROCESS_FOR_UB (tensorrt_llm.functional.AllReduceFusionOp attribute) layer_norm() (in module tensorrt_llm.functional) layer_quant_mode (tensorrt_llm.llmapi.QuantConfig property) layer_types (tensorrt_llm.runtime.ModelConfig attribute) layer_wise_benchmarks_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) LayerNorm (class in tensorrt_llm.layers.normalization) (tensorrt_llm.functional.LayerNormType attribute) LayerNormPositionType (class in tensorrt_llm.functional) LayerNormType (class in tensorrt_llm.functional) learned_absolute (tensorrt_llm.functional.PositionEmbeddingType attribute) length (tensorrt_llm.llmapi.CompletionOutput attribute) (tensorrt_llm.llmapi.CompletionOutput property) length_penalty (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) Linear (class in tensorrt_llm.layers.linear) linear (tensorrt_llm.functional.RotaryScalingType attribute) LinearActivation (class in tensorrt_llm.layers.mlp) LinearApproximateGELU (class in tensorrt_llm.layers.mlp) LinearBase (class in tensorrt_llm.layers.linear) LinearGEGLU (class in tensorrt_llm.layers.mlp) LinearGELU (class in tensorrt_llm.layers.mlp) LinearSwiGLU (class in tensorrt_llm.layers.mlp) ljust() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) llama3 (tensorrt_llm.functional.RotaryScalingType attribute) LLaMAConfig (class in tensorrt_llm.models) LLaMAForCausalLM (class in tensorrt_llm.models) LLaMAModel (class in tensorrt_llm.models) LlavaNextVisionConfig (class in tensorrt_llm.models) LlavaNextVisionWrapper (class in tensorrt_llm.models) LLM (class in tensorrt_llm.llmapi) llm_engine_dir (tensorrt_llm.runtime.MultimodalModelRunner property) llm_id (tensorrt_llm.llmapi.AsyncLLM property) (tensorrt_llm.llmapi.LLM attribute) (tensorrt_llm.llmapi.LLM property) (tensorrt_llm.llmapi.MultimodalEncoder property) LlmArgs (in module tensorrt_llm.llmapi) load() (tensorrt_llm.models.PretrainedModel method) (tensorrt_llm.models.SD3Transformer2DModel method) load_balancer (tensorrt_llm.llmapi.MoeConfig attribute) load_format (tensorrt_llm.llmapi.AutoDecodingConfig attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) load_test_audio() (tensorrt_llm.runtime.MultimodalModelRunner method) load_test_data() (tensorrt_llm.runtime.MultimodalModelRunner method) locate_accepted_draft_tokens() (tensorrt_llm.runtime.GenerationSession method) location (tensorrt_llm.functional.Tensor property) log() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) log_field_changes() (tensorrt_llm.plugin.PluginConfig class method) log_softmax() (in module tensorrt_llm.functional) log_two_model_deprecation_warning() (tensorrt_llm.llmapi.MTPDecodingConfig method) logits_processor (tensorrt_llm.llmapi.SamplingParams attribute) LogitsProcessor (class in tensorrt_llm.runtime) LogitsProcessorList (class in tensorrt_llm.runtime) logprobs (tensorrt_llm.llmapi.CompletionOutput attribute) (tensorrt_llm.llmapi.SamplingParams attribute) logprobs_diff (tensorrt_llm.llmapi.CompletionOutput attribute) (tensorrt_llm.llmapi.CompletionOutput property) logprobs_mode (tensorrt_llm.llmapi.SamplingParams attribute) long_rope (tensorrt_llm.functional.PositionEmbeddingType attribute) longrope (tensorrt_llm.functional.RotaryScalingType attribute) lookahead_config (tensorrt_llm.llmapi.SamplingParams attribute) LOOKAHEAD_DECODING (tensorrt_llm.models.SpeculativeDecodingMode attribute) LookaheadDecodingConfig (class in tensorrt_llm.llmapi) LookaheadDecodingConfig.Config (class in tensorrt_llm.llmapi) lora_ckpt_source (tensorrt_llm.llmapi.LoRARequest attribute) lora_config (tensorrt_llm.llmapi.BuildConfig attribute) (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) lora_int_id (tensorrt_llm.llmapi.LoRARequest attribute) lora_name (tensorrt_llm.llmapi.LoRARequest attribute) lora_path (tensorrt_llm.llmapi.LoRARequest attribute) lora_plugin (tensorrt_llm.runtime.ModelConfig attribute) lora_plugin() (in module tensorrt_llm.functional) lora_target_modules (tensorrt_llm.runtime.ModelConfig attribute) LoRARequest (class in tensorrt_llm.llmapi) low_latency_gemm() (in module tensorrt_llm.functional) low_latency_gemm_swiglu() (in module tensorrt_llm.functional) lower() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) LOWPRECISION (tensorrt_llm.functional.AllReduceStrategy attribute) lstrip() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) lt() (in module tensorrt_llm.functional) M make_causal_mask() (in module tensorrt_llm.layers.attention) maketrans() (tensorrt_llm.llmapi.BatchingType static method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy static method) (tensorrt_llm.llmapi.ContextChunkingPolicy static method) (tensorrt_llm.llmapi.QuantAlgo static method) mamba_conv1d() (in module tensorrt_llm.functional) mamba_conv1d_plugin (tensorrt_llm.runtime.ModelConfig attribute) mamba_ssm_cache_dtype (tensorrt_llm.llmapi.KvCacheConfig attribute) (tensorrt_llm.llmapi.QuantConfig attribute) mamba_ssm_philox_rounds (tensorrt_llm.llmapi.KvCacheConfig attribute) (tensorrt_llm.llmapi.QuantConfig attribute) mamba_ssm_stochastic_rounding (tensorrt_llm.llmapi.KvCacheConfig attribute) (tensorrt_llm.llmapi.QuantConfig attribute) mamba_state_cache_interval (tensorrt_llm.llmapi.KvCacheConfig attribute) MambaForCausalLM (class in tensorrt_llm.models) mapping (tensorrt_llm.runtime.GenerationSession attribute) (tensorrt_llm.runtime.ModelRunner property) mark_output() (tensorrt_llm.functional.Tensor method) mask_token_id (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) masked_scatter() (in module tensorrt_llm.functional) masked_select() (in module tensorrt_llm.functional) matmul() (in module tensorrt_llm.functional) max() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) max_attention_window (tensorrt_llm.llmapi.KvCacheConfig attribute) max_attention_window_size (tensorrt_llm.runtime.SamplingConfig attribute) max_batch_size (tensorrt_llm.llmapi.BuildConfig attribute) (tensorrt_llm.llmapi.CudaGraphConfig attribute) (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) (tensorrt_llm.runtime.ModelConfig attribute) max_beam_width (tensorrt_llm.llmapi.BuildConfig attribute) (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) (tensorrt_llm.runtime.ModelConfig attribute) max_cache_storage_gb (tensorrt_llm.llmapi.BuildCacheConfig attribute) max_concurrency (tensorrt_llm.llmapi.AutoDecodingConfig attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) max_draft_len (tensorrt_llm.llmapi.AutoDecodingConfig attribute) (tensorrt_llm.llmapi.BuildConfig attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) max_draft_tokens (tensorrt_llm.runtime.GenerationSession property) max_encoder_input_len (tensorrt_llm.llmapi.BuildConfig attribute) max_gpu_total_bytes (tensorrt_llm.llmapi.KvCacheConfig attribute) max_input_len (tensorrt_llm.llmapi.BuildConfig attribute) (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) max_matching_ngram_size (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) max_medusa_tokens (tensorrt_llm.runtime.ModelConfig attribute) max_new_tokens (tensorrt_llm.runtime.SamplingConfig attribute) max_ngram_size (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) max_non_leaves_per_layer (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) max_num_streams (tensorrt_llm.llmapi.TorchCompileConfig attribute) max_num_tokens (tensorrt_llm.llmapi.BuildConfig attribute) (tensorrt_llm.llmapi.MoeConfig attribute) (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) max_prompt_adapter_token (tensorrt_llm.llmapi.TrtLlmArgs attribute) max_prompt_embedding_table_size (tensorrt_llm.llmapi.BuildConfig attribute) (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) (tensorrt_llm.runtime.ModelRunner property) (tensorrt_llm.runtime.ModelRunnerCpp property) max_records (tensorrt_llm.llmapi.BuildCacheConfig attribute) max_seq_len (tensorrt_llm.llmapi.BuildConfig attribute) (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) max_sequence_length (tensorrt_llm.runtime.ModelRunner property) (tensorrt_llm.runtime.ModelRunnerCpp property) max_stats_len (tensorrt_llm.llmapi.TorchLlmArgs attribute) max_tokens (tensorrt_llm.llmapi.KvCacheConfig attribute) (tensorrt_llm.llmapi.SamplingParams attribute) max_tokens_in_buffer (tensorrt_llm.llmapi.CacheTransceiverConfig attribute) max_total_draft_tokens (tensorrt_llm.llmapi.AutoDecodingConfig attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) max_util_for_resume (tensorrt_llm.llmapi.KvCacheConfig attribute) MAX_UTILIZATION (tensorrt_llm.llmapi.CapacitySchedulerPolicy attribute) max_verification_set_size (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) max_window_size (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) maximum() (in module tensorrt_llm.functional) maybe_to_pybind() (tensorrt_llm.llmapi.CacheTransceiverConfig static method) (tensorrt_llm.llmapi.DynamicBatchConfig static method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig static method) (tensorrt_llm.llmapi.KvCacheConfig static method) (tensorrt_llm.llmapi.LookaheadDecodingConfig static method) (tensorrt_llm.llmapi.SchedulerConfig static method) mean() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) MEDUSA (tensorrt_llm.models.SpeculativeDecodingMode attribute) medusa_choices (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) medusa_decode_and_verify() (tensorrt_llm.runtime.GenerationSession method) medusa_paths (tensorrt_llm.runtime.GenerationSession attribute) medusa_position_offsets (tensorrt_llm.runtime.GenerationSession attribute) medusa_temperature (tensorrt_llm.runtime.GenerationSession attribute) medusa_topks (tensorrt_llm.runtime.GenerationSession attribute) medusa_tree_ids (tensorrt_llm.runtime.GenerationSession attribute) MedusaConfig (class in tensorrt_llm.models) MedusaDecodingConfig (class in tensorrt_llm.llmapi) MedusaDecodingConfig.Config (class in tensorrt_llm.llmapi) MedusaForCausalLm (class in tensorrt_llm.models) meshgrid2d() (in module tensorrt_llm.functional) metrics (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Output attribute) min() (in module tensorrt_llm.functional) MIN_LATENCY (tensorrt_llm.functional.AllReduceStrategy attribute) min_length (tensorrt_llm.runtime.SamplingConfig attribute) min_p (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) min_tokens (tensorrt_llm.llmapi.SamplingParams attribute) minimum() (in module tensorrt_llm.functional) mirror_pybind_enum() (tensorrt_llm.llmapi.CacheTransceiverConfig static method) (tensorrt_llm.llmapi.DynamicBatchConfig static method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig static method) (tensorrt_llm.llmapi.KvCacheConfig static method) (tensorrt_llm.llmapi.LookaheadDecodingConfig static method) (tensorrt_llm.llmapi.SchedulerConfig static method) mirror_pybind_fields() (tensorrt_llm.llmapi.CacheTransceiverConfig static method) (tensorrt_llm.llmapi.DynamicBatchConfig static method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig static method) (tensorrt_llm.llmapi.KvCacheConfig static method) (tensorrt_llm.llmapi.LookaheadDecodingConfig static method) (tensorrt_llm.llmapi.SchedulerConfig static method) Mish (class in tensorrt_llm.layers.activation) MIXED_PRECISION (tensorrt_llm.llmapi.QuantAlgo attribute) MLLaMAForCausalLM (class in tensorrt_llm.models) MLP (class in tensorrt_llm.layers.mlp) (tensorrt_llm.functional.MLPType attribute) MLPType (class in tensorrt_llm.functional) mm_encoder_only (tensorrt_llm.llmapi.TorchLlmArgs attribute) MNNVL (tensorrt_llm.functional.AllReduceStrategy attribute) MODEL trtllm-serve-mm_embedding_serve command line option trtllm-serve-serve command line option model (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) model_computed_fields (tensorrt_llm.llmapi.AttentionDpConfig attribute) (tensorrt_llm.llmapi.AutoDecodingConfig attribute) (tensorrt_llm.llmapi.BuildCacheConfig attribute) (tensorrt_llm.llmapi.BuildConfig attribute) (tensorrt_llm.llmapi.CacheTransceiverConfig attribute) (tensorrt_llm.llmapi.CalibConfig attribute) (tensorrt_llm.llmapi.CudaGraphConfig attribute) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig attribute) (tensorrt_llm.llmapi.DynamicBatchConfig attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig attribute) (tensorrt_llm.llmapi.KvCacheConfig attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.llmapi.MoeConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) (tensorrt_llm.llmapi.PrometheusMetricsConfig attribute) (tensorrt_llm.llmapi.QuantConfig attribute) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig attribute) (tensorrt_llm.llmapi.RocketSparseAttentionConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SAEnhancerConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) (tensorrt_llm.llmapi.SchedulerConfig attribute) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig attribute) (tensorrt_llm.llmapi.TorchCompileConfig attribute) (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) model_config (tensorrt_llm.llmapi.AttentionDpConfig attribute) (tensorrt_llm.llmapi.AutoDecodingConfig attribute) (tensorrt_llm.llmapi.BuildCacheConfig attribute) (tensorrt_llm.llmapi.BuildConfig attribute) (tensorrt_llm.llmapi.CacheTransceiverConfig attribute) (tensorrt_llm.llmapi.CalibConfig attribute) (tensorrt_llm.llmapi.CudaGraphConfig attribute) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig attribute) (tensorrt_llm.llmapi.DynamicBatchConfig attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig attribute) (tensorrt_llm.llmapi.KvCacheConfig attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.llmapi.MoeConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) (tensorrt_llm.llmapi.PrometheusMetricsConfig attribute) (tensorrt_llm.llmapi.QuantConfig attribute) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig attribute) (tensorrt_llm.llmapi.RocketSparseAttentionConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SAEnhancerConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) (tensorrt_llm.llmapi.SchedulerConfig attribute) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig attribute) (tensorrt_llm.llmapi.TorchCompileConfig attribute) (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) (tensorrt_llm.plugin.PluginConfig attribute) model_construct() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) model_copy() (tensorrt_llm.llmapi.AttentionDpConfig method) (tensorrt_llm.llmapi.AutoDecodingConfig method) (tensorrt_llm.llmapi.BuildCacheConfig method) (tensorrt_llm.llmapi.BuildConfig method) (tensorrt_llm.llmapi.CacheTransceiverConfig method) (tensorrt_llm.llmapi.CalibConfig method) (tensorrt_llm.llmapi.CudaGraphConfig method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig method) (tensorrt_llm.llmapi.DFlashDecodingConfig method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig method) (tensorrt_llm.llmapi.DynamicBatchConfig method) (tensorrt_llm.llmapi.Eagle3DecodingConfig method) (tensorrt_llm.llmapi.EagleDecodingConfig method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig method) (tensorrt_llm.llmapi.KvCacheConfig method) (tensorrt_llm.llmapi.LookaheadDecodingConfig method) (tensorrt_llm.llmapi.MedusaDecodingConfig method) (tensorrt_llm.llmapi.MoeConfig method) (tensorrt_llm.llmapi.MTPDecodingConfig method) (tensorrt_llm.llmapi.NGramDecodingConfig method) (tensorrt_llm.llmapi.PARDDecodingConfig method) (tensorrt_llm.llmapi.PrometheusMetricsConfig method) (tensorrt_llm.llmapi.QuantConfig method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig method) (tensorrt_llm.llmapi.SADecodingConfig method) (tensorrt_llm.llmapi.SAEnhancerConfig method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig method) (tensorrt_llm.llmapi.SchedulerConfig method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig method) (tensorrt_llm.llmapi.TorchCompileConfig method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig method) model_dump() (tensorrt_llm.llmapi.AttentionDpConfig method) (tensorrt_llm.llmapi.AutoDecodingConfig method) (tensorrt_llm.llmapi.BuildCacheConfig method) (tensorrt_llm.llmapi.BuildConfig method) (tensorrt_llm.llmapi.CacheTransceiverConfig method) (tensorrt_llm.llmapi.CalibConfig method) (tensorrt_llm.llmapi.CudaGraphConfig method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig method) (tensorrt_llm.llmapi.DFlashDecodingConfig method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig method) (tensorrt_llm.llmapi.DynamicBatchConfig method) (tensorrt_llm.llmapi.Eagle3DecodingConfig method) (tensorrt_llm.llmapi.EagleDecodingConfig method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig method) (tensorrt_llm.llmapi.KvCacheConfig method) (tensorrt_llm.llmapi.LookaheadDecodingConfig method) (tensorrt_llm.llmapi.MedusaDecodingConfig method) (tensorrt_llm.llmapi.MoeConfig method) (tensorrt_llm.llmapi.MTPDecodingConfig method) (tensorrt_llm.llmapi.NGramDecodingConfig method) (tensorrt_llm.llmapi.PARDDecodingConfig method) (tensorrt_llm.llmapi.PrometheusMetricsConfig method) (tensorrt_llm.llmapi.QuantConfig method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig method) (tensorrt_llm.llmapi.SADecodingConfig method) (tensorrt_llm.llmapi.SAEnhancerConfig method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig method) (tensorrt_llm.llmapi.SchedulerConfig method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig method) (tensorrt_llm.llmapi.TorchCompileConfig method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig method) model_dump_json() (tensorrt_llm.llmapi.AttentionDpConfig method) (tensorrt_llm.llmapi.AutoDecodingConfig method) (tensorrt_llm.llmapi.BuildCacheConfig method) (tensorrt_llm.llmapi.BuildConfig method) (tensorrt_llm.llmapi.CacheTransceiverConfig method) (tensorrt_llm.llmapi.CalibConfig method) (tensorrt_llm.llmapi.CudaGraphConfig method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig method) (tensorrt_llm.llmapi.DFlashDecodingConfig method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig method) (tensorrt_llm.llmapi.DynamicBatchConfig method) (tensorrt_llm.llmapi.Eagle3DecodingConfig method) (tensorrt_llm.llmapi.EagleDecodingConfig method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig method) (tensorrt_llm.llmapi.KvCacheConfig method) (tensorrt_llm.llmapi.LookaheadDecodingConfig method) (tensorrt_llm.llmapi.MedusaDecodingConfig method) (tensorrt_llm.llmapi.MoeConfig method) (tensorrt_llm.llmapi.MTPDecodingConfig method) (tensorrt_llm.llmapi.NGramDecodingConfig method) (tensorrt_llm.llmapi.PARDDecodingConfig method) (tensorrt_llm.llmapi.PrometheusMetricsConfig method) (tensorrt_llm.llmapi.QuantConfig method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig method) (tensorrt_llm.llmapi.SADecodingConfig method) (tensorrt_llm.llmapi.SAEnhancerConfig method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig method) (tensorrt_llm.llmapi.SchedulerConfig method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig method) (tensorrt_llm.llmapi.TorchCompileConfig method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig method) model_extra (tensorrt_llm.llmapi.AttentionDpConfig property) (tensorrt_llm.llmapi.AutoDecodingConfig property) (tensorrt_llm.llmapi.BuildCacheConfig property) (tensorrt_llm.llmapi.BuildConfig property) (tensorrt_llm.llmapi.CacheTransceiverConfig property) (tensorrt_llm.llmapi.CalibConfig property) (tensorrt_llm.llmapi.CudaGraphConfig property) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig property) (tensorrt_llm.llmapi.DFlashDecodingConfig property) (tensorrt_llm.llmapi.DraftTargetDecodingConfig property) (tensorrt_llm.llmapi.DynamicBatchConfig property) (tensorrt_llm.llmapi.Eagle3DecodingConfig property) (tensorrt_llm.llmapi.EagleDecodingConfig property) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig property) (tensorrt_llm.llmapi.KvCacheConfig property) (tensorrt_llm.llmapi.LookaheadDecodingConfig property) (tensorrt_llm.llmapi.MedusaDecodingConfig property) (tensorrt_llm.llmapi.MoeConfig property) (tensorrt_llm.llmapi.MTPDecodingConfig property) (tensorrt_llm.llmapi.NGramDecodingConfig property) (tensorrt_llm.llmapi.PARDDecodingConfig property) (tensorrt_llm.llmapi.PrometheusMetricsConfig property) (tensorrt_llm.llmapi.QuantConfig property) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig property) (tensorrt_llm.llmapi.RocketSparseAttentionConfig property) (tensorrt_llm.llmapi.SADecodingConfig property) (tensorrt_llm.llmapi.SAEnhancerConfig property) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig property) (tensorrt_llm.llmapi.SchedulerConfig property) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig property) (tensorrt_llm.llmapi.TorchCompileConfig property) (tensorrt_llm.llmapi.UserProvidedDecodingConfig property) model_fields (tensorrt_llm.llmapi.AttentionDpConfig attribute) (tensorrt_llm.llmapi.AutoDecodingConfig attribute) (tensorrt_llm.llmapi.BuildCacheConfig attribute) (tensorrt_llm.llmapi.BuildConfig attribute) (tensorrt_llm.llmapi.CacheTransceiverConfig attribute) (tensorrt_llm.llmapi.CalibConfig attribute) (tensorrt_llm.llmapi.CudaGraphConfig attribute) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig attribute) (tensorrt_llm.llmapi.DynamicBatchConfig attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig attribute) (tensorrt_llm.llmapi.KvCacheConfig attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.llmapi.MoeConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) (tensorrt_llm.llmapi.PrometheusMetricsConfig attribute) (tensorrt_llm.llmapi.QuantConfig attribute) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig attribute) (tensorrt_llm.llmapi.RocketSparseAttentionConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SAEnhancerConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) (tensorrt_llm.llmapi.SchedulerConfig attribute) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig attribute) (tensorrt_llm.llmapi.TorchCompileConfig attribute) (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) model_fields_set (tensorrt_llm.llmapi.AttentionDpConfig property) (tensorrt_llm.llmapi.AutoDecodingConfig property) (tensorrt_llm.llmapi.BuildCacheConfig property) (tensorrt_llm.llmapi.BuildConfig property) (tensorrt_llm.llmapi.CacheTransceiverConfig property) (tensorrt_llm.llmapi.CalibConfig property) (tensorrt_llm.llmapi.CudaGraphConfig property) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig property) (tensorrt_llm.llmapi.DFlashDecodingConfig property) (tensorrt_llm.llmapi.DraftTargetDecodingConfig property) (tensorrt_llm.llmapi.DynamicBatchConfig property) (tensorrt_llm.llmapi.Eagle3DecodingConfig property) (tensorrt_llm.llmapi.EagleDecodingConfig property) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig property) (tensorrt_llm.llmapi.KvCacheConfig property) (tensorrt_llm.llmapi.LookaheadDecodingConfig property) (tensorrt_llm.llmapi.MedusaDecodingConfig property) (tensorrt_llm.llmapi.MoeConfig property) (tensorrt_llm.llmapi.MTPDecodingConfig property) (tensorrt_llm.llmapi.NGramDecodingConfig property) (tensorrt_llm.llmapi.PARDDecodingConfig property) (tensorrt_llm.llmapi.PrometheusMetricsConfig property) (tensorrt_llm.llmapi.QuantConfig property) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig property) (tensorrt_llm.llmapi.RocketSparseAttentionConfig property) (tensorrt_llm.llmapi.SADecodingConfig property) (tensorrt_llm.llmapi.SAEnhancerConfig property) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig property) (tensorrt_llm.llmapi.SchedulerConfig property) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig property) (tensorrt_llm.llmapi.TorchCompileConfig property) (tensorrt_llm.llmapi.UserProvidedDecodingConfig property) model_format (tensorrt_llm.llmapi.TorchLlmArgs property) (tensorrt_llm.llmapi.TrtLlmArgs property) model_json_schema() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) model_kwargs (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) model_name (tensorrt_llm.runtime.ModelConfig attribute) model_parametrized_name() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) model_post_init() (tensorrt_llm.llmapi.AttentionDpConfig method) (tensorrt_llm.llmapi.AutoDecodingConfig method) (tensorrt_llm.llmapi.BuildCacheConfig method) (tensorrt_llm.llmapi.BuildConfig method) (tensorrt_llm.llmapi.CacheTransceiverConfig method) (tensorrt_llm.llmapi.CalibConfig method) (tensorrt_llm.llmapi.CudaGraphConfig method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig method) (tensorrt_llm.llmapi.DFlashDecodingConfig method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig method) (tensorrt_llm.llmapi.DynamicBatchConfig method) (tensorrt_llm.llmapi.Eagle3DecodingConfig method) (tensorrt_llm.llmapi.EagleDecodingConfig method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig method) (tensorrt_llm.llmapi.KvCacheConfig method) (tensorrt_llm.llmapi.LookaheadDecodingConfig method) (tensorrt_llm.llmapi.MedusaDecodingConfig method) (tensorrt_llm.llmapi.MoeConfig method) (tensorrt_llm.llmapi.MTPDecodingConfig method) (tensorrt_llm.llmapi.NGramDecodingConfig method) (tensorrt_llm.llmapi.PARDDecodingConfig method) (tensorrt_llm.llmapi.PrometheusMetricsConfig method) (tensorrt_llm.llmapi.QuantConfig method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig method) (tensorrt_llm.llmapi.SADecodingConfig method) (tensorrt_llm.llmapi.SAEnhancerConfig method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig method) (tensorrt_llm.llmapi.SchedulerConfig method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig method) (tensorrt_llm.llmapi.TorchCompileConfig method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig method) (tensorrt_llm.plugin.PluginConfig method) model_rebuild() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) model_validate() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) model_validate_json() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) model_validate_strings() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) ModelConfig (class in tensorrt_llm.runtime) ModelRunner (class in tensorrt_llm.runtime) ModelRunnerCpp (class in tensorrt_llm.runtime) module tensorrt_llm, [1], [2], [3], [4], [5] tensorrt_llm.functional tensorrt_llm.layers.activation tensorrt_llm.layers.attention tensorrt_llm.layers.cast tensorrt_llm.layers.conv tensorrt_llm.layers.embedding tensorrt_llm.layers.linear tensorrt_llm.layers.mlp tensorrt_llm.layers.normalization tensorrt_llm.layers.pooling tensorrt_llm.models tensorrt_llm.plugin tensorrt_llm.quantization tensorrt_llm.runtime modulo() (in module tensorrt_llm.functional) moe (tensorrt_llm.functional.SideStreamIDType attribute) moe_cluster_parallel_size (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) moe_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) moe_expert_parallel_size (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) MOE_FINALIZE_ALLREDUCE_RESIDUAL_RMS_NORM (tensorrt_llm.functional.AllReduceFusionOp attribute) moe_tensor_parallel_size (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) MoEAllReduceParams (class in tensorrt_llm.functional) MoeConfig (class in tensorrt_llm.llmapi) MoeConfig.Config (class in tensorrt_llm.llmapi) monitor_memory (tensorrt_llm.llmapi.BuildConfig attribute) mpi_session (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) MpiCommSession (class in tensorrt_llm.llmapi) MPTForCausalLM (class in tensorrt_llm.models) MPTModel (class in tensorrt_llm.models) mrope (tensorrt_llm.functional.PositionEmbeddingType attribute) (tensorrt_llm.functional.RotaryScalingType attribute) mrope_position_deltas_handle (tensorrt_llm.llmapi.DisaggregatedParams attribute) mrope_position_ids_handle (tensorrt_llm.llmapi.DisaggregatedParams attribute) MropeParams (class in tensorrt_llm.layers.attention) msg (tensorrt_llm.llmapi.TorchLlmArgs attribute), [1] (tensorrt_llm.llmapi.TrtLlmArgs attribute) mtp_eagle_one_model (tensorrt_llm.llmapi.MTPDecodingConfig attribute) MTPDecodingConfig (class in tensorrt_llm.llmapi) MTPDecodingConfig.Config (class in tensorrt_llm.llmapi) mul() (in module tensorrt_llm.functional) multi_block_mode (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig attribute) multimodal_embedding_handles (tensorrt_llm.llmapi.DisaggregatedParams attribute) multimodal_hashes (tensorrt_llm.llmapi.DisaggregatedParams attribute) MultimodalEncoder (class in tensorrt_llm.llmapi) MultimodalModelRunner (class in tensorrt_llm.runtime) multiply_and_lora() (tensorrt_llm.layers.linear.LinearBase method) multiply_collect() (tensorrt_llm.layers.linear.LinearBase method) (tensorrt_llm.layers.linear.RowLinear method) mx_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) N n (tensorrt_llm.llmapi.SamplingParams attribute) name (tensorrt_llm.functional.Tensor property) (tensorrt_llm.llmapi.LoRARequest property) (tensorrt_llm.runtime.TensorInfo attribute) NATIVE_QUANT_FLOW (tensorrt_llm.models.GemmaForCausalLM attribute) NCCL (tensorrt_llm.functional.AllReduceStrategy attribute) NCCL_SYMMETRIC (tensorrt_llm.functional.AllReduceStrategy attribute) ndim() (tensorrt_llm.functional.Tensor method) needs_separate_short_long_cuda_graphs() (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig method) network (tensorrt_llm.functional.Tensor property) next_medusa_input_ids() (tensorrt_llm.runtime.GenerationSession method) NGRAM (tensorrt_llm.models.SpeculativeDecodingMode attribute) NGramDecodingConfig (class in tensorrt_llm.llmapi) NGramDecodingConfig.Config (class in tensorrt_llm.llmapi) NO_QUANT (tensorrt_llm.llmapi.QuantAlgo attribute) no_repeat_ngram_size (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) non_gated_version() (in module tensorrt_llm.functional) NONE (tensorrt_llm.functional.AllReduceFusionOp attribute) none (tensorrt_llm.functional.RotaryScalingType attribute) NONE (tensorrt_llm.models.SpeculativeDecodingMode attribute) nonzero() (in module tensorrt_llm.functional) normalize_log_probs (tensorrt_llm.llmapi.TrtLlmArgs attribute) normalize_optional_fields_to_defaults() (tensorrt_llm.llmapi.TorchLlmArgs method) (tensorrt_llm.llmapi.TrtLlmArgs method) not_op() (in module tensorrt_llm.functional) num_beams (tensorrt_llm.runtime.SamplingConfig attribute) num_capture_layers (tensorrt_llm.llmapi.Eagle3DecodingConfig property) (tensorrt_llm.llmapi.EagleDecodingConfig property) (tensorrt_llm.llmapi.MTPDecodingConfig property) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig property) num_capture_layers() (tensorrt_llm.llmapi.AutoDecodingConfig method) (tensorrt_llm.llmapi.DFlashDecodingConfig method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig method) (tensorrt_llm.llmapi.LookaheadDecodingConfig method) (tensorrt_llm.llmapi.MedusaDecodingConfig method) (tensorrt_llm.llmapi.NGramDecodingConfig method) (tensorrt_llm.llmapi.PARDDecodingConfig method) (tensorrt_llm.llmapi.SADecodingConfig method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig method) num_draft_tokens (tensorrt_llm.runtime.GenerationSession attribute) num_eagle_layers (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) num_heads (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) (tensorrt_llm.runtime.ModelRunner property) (tensorrt_llm.runtime.ModelRunnerCpp property) num_kv_heads (tensorrt_llm.runtime.ModelConfig attribute) num_kv_heads_per_cross_attn_layer (tensorrt_llm.runtime.ModelConfig attribute) num_kv_heads_per_layer (tensorrt_llm.runtime.ModelConfig attribute) num_layers (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) (tensorrt_llm.runtime.ModelRunner property) (tensorrt_llm.runtime.ModelRunnerCpp property) num_medusa_heads (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) num_nextn_predict_layers (tensorrt_llm.llmapi.MTPDecodingConfig attribute) num_postprocess_workers (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) num_return_sequences (tensorrt_llm.runtime.SamplingConfig attribute) numel() (tensorrt_llm.runtime.TensorInfo method) numerator (tensorrt_llm.llmapi.DisaggScheduleStyle attribute) NVFP4 (tensorrt_llm.llmapi.QuantAlgo attribute) NVFP4_ARC (tensorrt_llm.llmapi.QuantAlgo attribute) NVFP4_AWQ (tensorrt_llm.llmapi.QuantAlgo attribute) nvfp4_gemm_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) O ONESHOT (tensorrt_llm.functional.AllReduceStrategy attribute) op_and() (in module tensorrt_llm.functional) op_or() (in module tensorrt_llm.functional) op_xor() (in module tensorrt_llm.functional) opaque_state (tensorrt_llm.llmapi.DisaggregatedParams attribute) opt_batch_size (tensorrt_llm.llmapi.BuildConfig attribute) opt_num_tokens (tensorrt_llm.llmapi.BuildConfig attribute) OPTForCausalLM (class in tensorrt_llm.models) OPTModel (class in tensorrt_llm.models) orchestrator_type (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) otlp_traces_endpoint (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) outer() (in module tensorrt_llm.functional) output_cum_log_probs (tensorrt_llm.runtime.SamplingConfig attribute) output_directory (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) output_log_probs (tensorrt_llm.runtime.SamplingConfig attribute) output_sequence_lengths (tensorrt_llm.runtime.SamplingConfig attribute) output_timing_cache (tensorrt_llm.llmapi.BuildConfig attribute) outputs (tensorrt_llm.llmapi.RequestOutput attribute) (tensorrt_llm.llmapi.RequestOutput property) P pad() (in module tensorrt_llm.functional) pad_id (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) padding (tensorrt_llm.functional.AttentionMaskType attribute) page_size (tensorrt_llm.llmapi.RocketSparseAttentionConfig attribute) paged_kv_cache (tensorrt_llm.runtime.GenerationSession property) paged_state (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) parallel_config (tensorrt_llm.llmapi.TorchLlmArgs property) (tensorrt_llm.llmapi.TrtLlmArgs property) params_imply_greedy_decoding() (tensorrt_llm.llmapi.SamplingParams static method) PARDDecodingConfig (class in tensorrt_llm.llmapi) PARDDecodingConfig.Config (class in tensorrt_llm.llmapi) parse_file() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) parse_obj() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) parse_raw() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) partition() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) path (tensorrt_llm.llmapi.LoRARequest property) pause_generation() (tensorrt_llm.llmapi.AsyncLLM method) peft_cache_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) perf_metrics_max_requests (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) permute() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) Phi3ForCausalLM (class in tensorrt_llm.models) Phi3Model (class in tensorrt_llm.models) PhiForCausalLM (class in tensorrt_llm.models) PhiModel (class in tensorrt_llm.models) pipeline_parallel_size (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) PixArtAlphaTextProjection (class in tensorrt_llm.layers.embedding) plugin_config (tensorrt_llm.llmapi.BuildConfig attribute) PluginConfig (class in tensorrt_llm.plugin) policy_args (tensorrt_llm.llmapi.ReorderRequestPolicyConfig attribute) policy_name (tensorrt_llm.llmapi.ReorderRequestPolicyConfig attribute) PositionEmbeddingType (class in tensorrt_llm.functional) post_layernorm (tensorrt_llm.functional.LayerNormPositionType attribute) posterior_threshold (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) postproc_params (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Input attribute) postprocess() (tensorrt_llm.layers.attention.Attention method) (tensorrt_llm.layers.attention.DeepseekV2Attention method) (tensorrt_llm.layers.embedding.Embedding method) (tensorrt_llm.layers.linear.Linear method) postprocess_tokenizer_dir (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) pow() (in module tensorrt_llm.functional) pp_communicate_final_output_ids() (tensorrt_llm.runtime.GenerationSession method) pp_communicate_new_tokens() (tensorrt_llm.runtime.GenerationSession method) pp_partition (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) pre_layernorm (tensorrt_llm.functional.LayerNormPositionType attribute) pre_quant_scale (tensorrt_llm.llmapi.QuantConfig attribute) precompute_relative_attention_bias() (tensorrt_llm.models.DecoderModel method) (tensorrt_llm.models.EncoderModel method) (tensorrt_llm.models.WhisperEncoder method) prepare_inputs() (tensorrt_llm.models.ChatGLMForCausalLM method) (tensorrt_llm.models.DecoderModel method) (tensorrt_llm.models.DiT method) (tensorrt_llm.models.EagleForCausalLM method) (tensorrt_llm.models.EncoderModel method) (tensorrt_llm.models.LlavaNextVisionWrapper method) (tensorrt_llm.models.MambaForCausalLM method) (tensorrt_llm.models.MLLaMAForCausalLM method) (tensorrt_llm.models.PretrainedModel method) (tensorrt_llm.models.RecurrentGemmaForCausalLM method) (tensorrt_llm.models.SD3Transformer2DModel method) (tensorrt_llm.models.WhisperEncoder method) prepare_position_ids_for_cogvlm() (tensorrt_llm.runtime.MultimodalModelRunner method) prepare_recurrent_inputs() (tensorrt_llm.models.RecurrentGemmaForCausalLM method) preprocess() (tensorrt_llm.llmapi.AsyncLLM method) (tensorrt_llm.llmapi.LLM method) (tensorrt_llm.llmapi.MultimodalEncoder method) (tensorrt_llm.runtime.MultimodalModelRunner method) presence_penalty (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) PretrainedConfig (class in tensorrt_llm.models) PretrainedModel (class in tensorrt_llm.models) print_iter_log (tensorrt_llm.llmapi.TorchLlmArgs attribute) priority (tensorrt_llm.llmapi.KvCacheRetentionConfig.TokenRangeRetentionConfig property) process_input() (tensorrt_llm.runtime.EncDecModelRunner method) process_logits_including_draft() (tensorrt_llm.runtime.GenerationSession method) prod() (in module tensorrt_llm.functional) profiler (tensorrt_llm.runtime.GenerationSession property) profiling_verbosity (tensorrt_llm.llmapi.BuildConfig attribute) prometheus_metrics_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) PrometheusMetricsConfig (class in tensorrt_llm.llmapi) PrometheusMetricsConfig.Config (class in tensorrt_llm.llmapi) prompt (tensorrt_llm.llmapi.RequestOutput attribute) (tensorrt_llm.llmapi.RequestOutput property) prompt_budget (tensorrt_llm.llmapi.RocketSparseAttentionConfig attribute) prompt_ignore_length (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) prompt_logprobs (tensorrt_llm.llmapi.CompletionOutput attribute) (tensorrt_llm.llmapi.SamplingParams attribute) prompt_token_ids (tensorrt_llm.llmapi.RequestOutput attribute) (tensorrt_llm.llmapi.RequestOutput property) PromptTuningEmbedding (class in tensorrt_llm.layers.embedding) ptuning_setup() (tensorrt_llm.runtime.MultimodalModelRunner method) ptuning_setup_fuyu() (tensorrt_llm.runtime.MultimodalModelRunner method) ptuning_setup_llava_next() (tensorrt_llm.runtime.MultimodalModelRunner method) ptuning_setup_phi3() (tensorrt_llm.runtime.MultimodalModelRunner method) ptuning_setup_pixtral() (tensorrt_llm.runtime.MultimodalModelRunner method) pybind_equals() (tensorrt_llm.llmapi.CacheTransceiverConfig static method) (tensorrt_llm.llmapi.DynamicBatchConfig static method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig static method) (tensorrt_llm.llmapi.KvCacheConfig static method) (tensorrt_llm.llmapi.LookaheadDecodingConfig static method) (tensorrt_llm.llmapi.SchedulerConfig static method) python_e2e (tensorrt_llm.runtime.MultimodalModelRunner property) Q q_split_threshold (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) quant_algo (tensorrt_llm.llmapi.QuantConfig attribute) (tensorrt_llm.models.PretrainedConfig property) quant_config (tensorrt_llm.llmapi.TorchLlmArgs property) (tensorrt_llm.llmapi.TrtLlmArgs attribute) quant_mode (tensorrt_llm.llmapi.QuantConfig property) (tensorrt_llm.models.PretrainedConfig property) (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) QuantAlgo (class in tensorrt_llm.llmapi) (class in tensorrt_llm.quantization) QuantConfig (class in tensorrt_llm.llmapi) QuantConfig.Config (class in tensorrt_llm.llmapi) quantize() (tensorrt_llm.models.BaichuanForCausalLM class method) (tensorrt_llm.models.ChatGLMForCausalLM class method) (tensorrt_llm.models.CogVLMForCausalLM class method) (tensorrt_llm.models.GemmaForCausalLM class method) (tensorrt_llm.models.GPTForCausalLM class method) (tensorrt_llm.models.LLaMAForCausalLM class method) (tensorrt_llm.models.PretrainedModel class method) quantize_and_export() (in module tensorrt_llm.quantization) QuantMode (class in tensorrt_llm.quantization) quick_gelu() (in module tensorrt_llm.functional) QWenForCausalLMGenerationSession (class in tensorrt_llm.runtime) R rand() (in module tensorrt_llm.functional) random_seed (tensorrt_llm.llmapi.CalibConfig attribute) (tensorrt_llm.runtime.SamplingConfig attribute) rank() (tensorrt_llm.functional.Tensor method) ray_placement_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) ray_worker_extension_cls (tensorrt_llm.llmapi.TorchLlmArgs attribute) ray_worker_nsight_options (tensorrt_llm.llmapi.TorchLlmArgs attribute) real (tensorrt_llm.llmapi.DisaggScheduleStyle attribute) rearrange() (in module tensorrt_llm.functional) reasoning_parser (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) record_stats() (tensorrt_llm.llmapi.RequestOutput method) RecurrentGemmaForCausalLM (class in tensorrt_llm.models) recv() (in module tensorrt_llm.functional) redrafter_draft_len_per_beam (tensorrt_llm.runtime.ModelConfig attribute) redrafter_num_beams (tensorrt_llm.runtime.ModelConfig attribute) ReDrafterForLLaMALM (class in tensorrt_llm.models) ReDrafterForQWenLM (class in tensorrt_llm.models) reduce() (in module tensorrt_llm.functional) reduce_scatter() (in module tensorrt_llm.functional) regex (tensorrt_llm.llmapi.GuidedDecodingParams attribute) relative (tensorrt_llm.functional.PositionEmbeddingType attribute) relaxed_delta (tensorrt_llm.llmapi.MTPDecodingConfig attribute) relaxed_topk (tensorrt_llm.llmapi.MTPDecodingConfig attribute) release() (tensorrt_llm.llmapi.AsyncLLM method) (tensorrt_llm.models.PretrainedModel method) relu() (in module tensorrt_llm.functional) remove_input_padding (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) (tensorrt_llm.runtime.ModelRunner property) (tensorrt_llm.runtime.ModelRunnerCpp property) removeprefix() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) removesuffix() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) reorder_kv_cache_for_beam_search() (tensorrt_llm.runtime.GenerationSession method) reorder_policy_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) ReorderRequestPolicyConfig (class in tensorrt_llm.llmapi) ReorderRequestPolicyConfig.Config (class in tensorrt_llm.llmapi) repeat() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) repeat_interleave() (in module tensorrt_llm.functional) repetition_penalty (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) replace() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) replace_all_uses_with() (tensorrt_llm.functional.Tensor method) request_decode_time_buckets (tensorrt_llm.llmapi.PrometheusMetricsConfig attribute) request_id (tensorrt_llm.llmapi.RequestOutput attribute) (tensorrt_llm.llmapi.RequestOutput property) request_inference_time_buckets (tensorrt_llm.llmapi.PrometheusMetricsConfig attribute) request_perf_metrics (tensorrt_llm.llmapi.CompletionOutput attribute) (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Output attribute) request_prefill_time_buckets (tensorrt_llm.llmapi.PrometheusMetricsConfig attribute) request_queue_time_buckets (tensorrt_llm.llmapi.PrometheusMetricsConfig attribute) request_stats_max_iterations (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) request_type (tensorrt_llm.llmapi.DisaggregatedParams attribute) RequestError (class in tensorrt_llm.llmapi) RequestOutput (class in tensorrt_llm.llmapi) RequestOutput.PostprocWorker (class in tensorrt_llm.llmapi) RequestOutput.PostprocWorker.Input (class in tensorrt_llm.llmapi) RequestOutput.PostprocWorker.Output (class in tensorrt_llm.llmapi) res (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Output attribute) RESIDUAL_RMS_NORM (tensorrt_llm.functional.AllReduceFusionOp attribute) RESIDUAL_RMS_NORM_OUT_QUANT_FP8 (tensorrt_llm.functional.AllReduceFusionOp attribute) RESIDUAL_RMS_NORM_OUT_QUANT_NVFP4 (tensorrt_llm.functional.AllReduceFusionOp attribute) RESIDUAL_RMS_NORM_QUANT_FP8 (tensorrt_llm.functional.AllReduceFusionOp attribute) RESIDUAL_RMS_NORM_QUANT_NVFP4 (tensorrt_llm.functional.AllReduceFusionOp attribute) RESIDUAL_RMS_PREPOST_NORM (tensorrt_llm.functional.AllReduceFusionOp attribute) resolve_for_target_sparsity() (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig method) resource_manager (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) result() (tensorrt_llm.llmapi.RequestOutput method) resume() (tensorrt_llm.llmapi.AsyncLLM method) resume_generation() (tensorrt_llm.llmapi.AsyncLLM method) return_context_logits (tensorrt_llm.llmapi.SamplingParams attribute) return_dict (tensorrt_llm.runtime.SamplingConfig attribute) return_encoder_output (tensorrt_llm.llmapi.SamplingParams attribute) return_generation_logits (tensorrt_llm.llmapi.SamplingParams attribute) return_perf_metrics (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) revision (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) rfind() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) rg_lru() (in module tensorrt_llm.functional) rindex() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) rjust() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) RMS_NORM (tensorrt_llm.functional.AllReduceFusionOp attribute) rms_norm() (in module tensorrt_llm.functional) RmsNorm (class in tensorrt_llm.layers.normalization) (tensorrt_llm.functional.LayerNormType attribute) rnn_conv_dim_size (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) rnn_head_size (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) rnn_hidden_size (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) RobertaForQuestionAnswering (in module tensorrt_llm.models) RobertaForSequenceClassification (in module tensorrt_llm.models) RobertaModel (in module tensorrt_llm.models) RocketSparseAttentionConfig (class in tensorrt_llm.llmapi) RocketSparseAttentionConfig.Config (class in tensorrt_llm.llmapi) rope_gpt_neox (tensorrt_llm.functional.PositionEmbeddingType attribute) rope_gptj (tensorrt_llm.functional.PositionEmbeddingType attribute) RopeEmbeddingUtils (class in tensorrt_llm.functional) RotaryScalingType (class in tensorrt_llm.functional) rotate_every_two() (tensorrt_llm.functional.RopeEmbeddingUtils static method) rotate_half() (tensorrt_llm.functional.RopeEmbeddingUtils static method) round() (in module tensorrt_llm.functional) RowLinear (class in tensorrt_llm.layers.linear) rpartition() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) rsp (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Input attribute) rsplit() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) rstrip() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) run() (tensorrt_llm.runtime.MultimodalModelRunner method) (tensorrt_llm.runtime.Session method) runtime (tensorrt_llm.runtime.GenerationSession attribute) (tensorrt_llm.runtime.Session property) S sa_config (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) SADecodingConfig (class in tensorrt_llm.llmapi) SADecodingConfig.Config (class in tensorrt_llm.llmapi) SAEnhancerConfig (class in tensorrt_llm.llmapi) SAEnhancerConfig.Config (class in tensorrt_llm.llmapi) sampler_force_async_worker (tensorrt_llm.llmapi.TorchLlmArgs attribute) sampler_type (tensorrt_llm.llmapi.TorchLlmArgs attribute) sampling_params (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Input attribute) SamplingConfig (class in tensorrt_llm.runtime) SamplingParams (class in tensorrt_llm.llmapi) save_checkpoint() (tensorrt_llm.models.LlavaNextVisionWrapper method) (tensorrt_llm.models.PretrainedModel method) SAVE_HIDDEN_STATES (tensorrt_llm.models.SpeculativeDecodingMode attribute) SaveHiddenStatesDecodingConfig (class in tensorrt_llm.llmapi) SaveHiddenStatesDecodingConfig.Config (class in tensorrt_llm.llmapi) scatter() (in module tensorrt_llm.functional) scatter_nd() (in module tensorrt_llm.functional) schedule_style (tensorrt_llm.llmapi.DisaggregatedParams attribute) scheduler_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) SchedulerConfig (class in tensorrt_llm.llmapi) SchedulerConfig.Config (class in tensorrt_llm.llmapi) SchedulingParams (class in tensorrt_llm.llmapi) schema() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) schema_json() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) SD35AdaLayerNormZeroX (class in tensorrt_llm.layers.normalization) SD3PatchEmbed (class in tensorrt_llm.layers.embedding) SD3Transformer2DModel (class in tensorrt_llm.models) secondary_offload_min_priority (tensorrt_llm.llmapi.KvCacheConfig attribute) seed (tensorrt_llm.llmapi.SamplingParams attribute) select() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) selective_scan() (in module tensorrt_llm.functional) send() (in module tensorrt_llm.functional) seq_len_threshold (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) (tensorrt_llm.llmapi.RocketSparseAttentionConfig attribute) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig attribute) serialize_engine() (tensorrt_llm.runtime.ModelRunner method) Session (class in tensorrt_llm.runtime) set_attn_processor() (tensorrt_llm.models.SD3Transformer2DModel method) set_default_cache_root() (tensorrt_llm.llmapi.BuildCacheConfig method) set_default_capture_num_tokens() (tensorrt_llm.llmapi.TorchCompileConfig method) set_if_not_exist() (tensorrt_llm.models.PretrainedConfig method) set_max_total_draft_tokens() (tensorrt_llm.llmapi.AutoDecodingConfig method) (tensorrt_llm.llmapi.DFlashDecodingConfig method) (tensorrt_llm.llmapi.LookaheadDecodingConfig method) (tensorrt_llm.llmapi.MedusaDecodingConfig method) (tensorrt_llm.llmapi.MTPDecodingConfig method) (tensorrt_llm.llmapi.PARDDecodingConfig method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig method) set_model_format() (tensorrt_llm.llmapi.TorchLlmArgs method) set_rank() (tensorrt_llm.models.PretrainedConfig method) set_rel_attn_table() (tensorrt_llm.layers.attention.Attention method) set_shapes() (tensorrt_llm.runtime.Session method) setup() (tensorrt_llm.runtime.GenerationSession method) setup_async() (tensorrt_llm.llmapi.AsyncLLM method) setup_embedding_parallel_mode() (tensorrt_llm.llmapi.TrtLlmArgs method) setup_fake_prompts() (tensorrt_llm.runtime.MultimodalModelRunner method) setup_fake_prompts_qwen2vl() (tensorrt_llm.runtime.MultimodalModelRunner method) setup_fake_prompts_vila() (tensorrt_llm.runtime.MultimodalModelRunner method) setup_inputs() (tensorrt_llm.runtime.MultimodalModelRunner method) shape (tensorrt_llm.functional.Tensor property) (tensorrt_llm.runtime.TensorInfo attribute) shape() (in module tensorrt_llm.functional) should_abort (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Output attribute) shutdown() (tensorrt_llm.llmapi.AsyncLLM method) (tensorrt_llm.llmapi.LLM method) (tensorrt_llm.llmapi.MpiCommSession method) (tensorrt_llm.llmapi.MultimodalEncoder method) shutdown_abort() (tensorrt_llm.llmapi.MpiCommSession method) SideStreamIDType (class in tensorrt_llm.functional) sigmoid() (in module tensorrt_llm.functional) silu() (in module tensorrt_llm.functional) sin() (in module tensorrt_llm.functional) sink_token_length (tensorrt_llm.llmapi.KvCacheConfig attribute) (tensorrt_llm.runtime.SamplingConfig attribute) size (tensorrt_llm.functional.SliceInputType attribute) size() (tensorrt_llm.functional.Tensor method) skip_cross_attn_blocks (tensorrt_llm.runtime.ModelConfig attribute) skip_cross_kv (tensorrt_llm.runtime.ModelConfig attribute) skip_indexer_for_short_seqs (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) skip_special_tokens (tensorrt_llm.llmapi.SamplingParams attribute) skip_tokenizer_init (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) SkipSoftmaxAttentionConfig (class in tensorrt_llm.llmapi) SkipSoftmaxAttentionConfig.Config (class in tensorrt_llm.llmapi) sleep_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) slice() (in module tensorrt_llm.functional) SliceInputType (class in tensorrt_llm.functional) sliding_window_causal (tensorrt_llm.functional.AttentionMaskType attribute) smoothquant_val (tensorrt_llm.llmapi.QuantConfig attribute) softmax() (in module tensorrt_llm.functional) softplus() (in module tensorrt_llm.functional) spaces_between_special_tokens (tensorrt_llm.llmapi.SamplingParams attribute) sparse_attention_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) spec_dec_mode (tensorrt_llm.llmapi.AutoDecodingConfig property) (tensorrt_llm.llmapi.DFlashDecodingConfig property) (tensorrt_llm.llmapi.DraftTargetDecodingConfig property) (tensorrt_llm.llmapi.Eagle3DecodingConfig property) (tensorrt_llm.llmapi.EagleDecodingConfig property) (tensorrt_llm.llmapi.LookaheadDecodingConfig property) (tensorrt_llm.llmapi.MedusaDecodingConfig property) (tensorrt_llm.llmapi.MTPDecodingConfig property) (tensorrt_llm.llmapi.NGramDecodingConfig property) (tensorrt_llm.llmapi.PARDDecodingConfig property) (tensorrt_llm.llmapi.SADecodingConfig property) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig property) (tensorrt_llm.llmapi.UserProvidedDecodingConfig property) SpecDecodingParams (class in tensorrt_llm.layers.attention) speculative_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) speculative_decoding_mode (tensorrt_llm.llmapi.BuildConfig attribute) speculative_model (tensorrt_llm.llmapi.AutoDecodingConfig attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) (tensorrt_llm.llmapi.TorchLlmArgs property) (tensorrt_llm.llmapi.TrtLlmArgs property) (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) SpeculativeDecodingMode (class in tensorrt_llm.models) split() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) split_prompt_by_images() (tensorrt_llm.runtime.MultimodalModelRunner method) splitlines() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) sqrt() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) squared_relu() (in module tensorrt_llm.functional) squeeze() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) (tensorrt_llm.runtime.TensorInfo method) stack() (in module tensorrt_llm.functional) start (tensorrt_llm.functional.SliceInputType attribute) start() (tensorrt_llm.llmapi.RequestOutput.PostprocWorker method) startswith() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) state_dtype (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) state_size (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) STATIC (tensorrt_llm.llmapi.BatchingType attribute) STATIC_BATCH (tensorrt_llm.llmapi.CapacitySchedulerPolicy attribute) step() (tensorrt_llm.runtime.KVCacheManager method) stop (tensorrt_llm.llmapi.SamplingParams attribute) stop_reason (tensorrt_llm.llmapi.CompletionOutput attribute) stop_token_ids (tensorrt_llm.llmapi.SamplingParams attribute) stop_words_list (tensorrt_llm.runtime.SamplingConfig attribute) StoppingCriteria (class in tensorrt_llm.runtime) StoppingCriteriaList (class in tensorrt_llm.runtime) stream_interval (tensorrt_llm.llmapi.TorchLlmArgs attribute) streaming (tensorrt_llm.llmapi.RequestOutput.PostprocWorker.Input attribute) stride (tensorrt_llm.functional.SliceInputType attribute) strip() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) strongly_typed (tensorrt_llm.llmapi.BuildConfig attribute) structural_tag (tensorrt_llm.llmapi.GuidedDecodingParams attribute) sub() (in module tensorrt_llm.functional) submit() (tensorrt_llm.llmapi.MpiCommSession method) submit_sync() (tensorrt_llm.llmapi.MpiCommSession method) sum() (in module tensorrt_llm.functional) supports_backend() (tensorrt_llm.llmapi.AutoDecodingConfig method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig method) (tensorrt_llm.llmapi.DFlashDecodingConfig method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig method) (tensorrt_llm.llmapi.Eagle3DecodingConfig method) (tensorrt_llm.llmapi.EagleDecodingConfig method) (tensorrt_llm.llmapi.LookaheadDecodingConfig method) (tensorrt_llm.llmapi.MedusaDecodingConfig method) (tensorrt_llm.llmapi.MTPDecodingConfig method) (tensorrt_llm.llmapi.NGramDecodingConfig method) (tensorrt_llm.llmapi.PARDDecodingConfig method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig method) (tensorrt_llm.llmapi.SADecodingConfig method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig method) swapcase() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) swiglu() (in module tensorrt_llm.functional) SYMM_MEM (tensorrt_llm.functional.AllReduceStrategy attribute) sync_quant_config_with_kv_cache_config_dtype() (tensorrt_llm.llmapi.TorchLlmArgs method) T tanh() (in module tensorrt_llm.functional) target_layer_ids (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) target_sparsity (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig attribute) target_sparsity_decode (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig property) target_sparsity_prefill (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig property) telemetry_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) temperature (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) Tensor (class in tensorrt_llm.functional) tensor_parallel_size (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) TensorInfo (class in tensorrt_llm.runtime) tensorrt_llm module, [1], [2], [3], [4], [5] tensorrt_llm.functional module tensorrt_llm.layers.activation module tensorrt_llm.layers.attention module tensorrt_llm.layers.cast module tensorrt_llm.layers.conv module tensorrt_llm.layers.embedding module tensorrt_llm.layers.linear module tensorrt_llm.layers.mlp module tensorrt_llm.layers.normalization module tensorrt_llm.layers.pooling module tensorrt_llm.models module tensorrt_llm.plugin module tensorrt_llm.quantization module tensorrt_llm.runtime module text (tensorrt_llm.llmapi.CompletionOutput attribute) text_diff (tensorrt_llm.llmapi.CompletionOutput attribute) (tensorrt_llm.llmapi.CompletionOutput property) threshold (tensorrt_llm.llmapi.SAEnhancerConfig attribute) threshold_scale_factor (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig attribute) threshold_scale_factor_decode (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig property) threshold_scale_factor_prefill (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig property) time_breakdown_metrics (tensorrt_llm.llmapi.RequestOutput attribute) time_per_output_token_buckets (tensorrt_llm.llmapi.PrometheusMetricsConfig attribute) time_to_first_token_buckets (tensorrt_llm.llmapi.PrometheusMetricsConfig attribute) timeout_iters (tensorrt_llm.llmapi.AttentionDpConfig attribute) TimestepEmbedding (class in tensorrt_llm.layers.embedding) Timesteps (class in tensorrt_llm.layers.embedding) title() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) to_bytes() (tensorrt_llm.llmapi.DisaggScheduleStyle method) to_dict() (tensorrt_llm.models.ChatGLMConfig method) (tensorrt_llm.models.CogVLMConfig method) (tensorrt_llm.models.DbrxConfig method) (tensorrt_llm.models.FalconConfig method) (tensorrt_llm.models.GemmaConfig method) (tensorrt_llm.models.GPTConfig method) (tensorrt_llm.models.GPTJConfig method) (tensorrt_llm.models.LLaMAConfig method) (tensorrt_llm.models.MedusaConfig method) (tensorrt_llm.models.PretrainedConfig method) to_json_file() (tensorrt_llm.models.PretrainedConfig method) to_layer_quant_config() (tensorrt_llm.models.PretrainedConfig method) to_legacy_setting() (tensorrt_llm.plugin.PluginConfig method) token_drop() (tensorrt_llm.layers.embedding.LabelEmbedding method) token_end (tensorrt_llm.llmapi.KvCacheRetentionConfig.TokenRangeRetentionConfig property) token_ids (tensorrt_llm.llmapi.CompletionOutput attribute) token_ids_diff (tensorrt_llm.llmapi.CompletionOutput attribute) (tensorrt_llm.llmapi.CompletionOutput property) token_range_retention_configs (tensorrt_llm.llmapi.KvCacheRetentionConfig property) token_start (tensorrt_llm.llmapi.KvCacheRetentionConfig.TokenRangeRetentionConfig property) tokenizer (tensorrt_llm.llmapi.AsyncLLM property) (tensorrt_llm.llmapi.LLM attribute) (tensorrt_llm.llmapi.LLM property) (tensorrt_llm.llmapi.MultimodalEncoder property) (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) tokenizer_image_token() (tensorrt_llm.runtime.MultimodalModelRunner static method) tokenizer_max_seq_length (tensorrt_llm.llmapi.CalibConfig attribute) tokenizer_mode (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) tokenizer_revision (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) tokens_per_block (tensorrt_llm.llmapi.KvCacheConfig attribute) (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) tokens_per_gen_step (tensorrt_llm.llmapi.AutoDecodingConfig property) (tensorrt_llm.llmapi.DFlashDecodingConfig property) (tensorrt_llm.llmapi.DraftTargetDecodingConfig property) (tensorrt_llm.llmapi.Eagle3DecodingConfig property) (tensorrt_llm.llmapi.EagleDecodingConfig property) (tensorrt_llm.llmapi.LookaheadDecodingConfig property) (tensorrt_llm.llmapi.MedusaDecodingConfig property) (tensorrt_llm.llmapi.MTPDecodingConfig property) (tensorrt_llm.llmapi.NGramDecodingConfig property) (tensorrt_llm.llmapi.PARDDecodingConfig property) (tensorrt_llm.llmapi.SADecodingConfig property) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig property) (tensorrt_llm.llmapi.UserProvidedDecodingConfig property) top_k (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) top_p (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) top_p_decay (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) top_p_min (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) top_p_reset_ids (tensorrt_llm.llmapi.SamplingParams attribute) (tensorrt_llm.runtime.SamplingConfig attribute) topk (tensorrt_llm.llmapi.RocketSparseAttentionConfig attribute) topk() (in module tensorrt_llm.functional) topr (tensorrt_llm.llmapi.RocketSparseAttentionConfig attribute) torch_compile_config (tensorrt_llm.llmapi.TorchLlmArgs attribute) TorchCompileConfig (class in tensorrt_llm.llmapi) TorchCompileConfig.Config (class in tensorrt_llm.llmapi) TorchLlmArgs (class in tensorrt_llm.llmapi) TorchLlmArgs.Config (class in tensorrt_llm.llmapi) tp_split_dim() (tensorrt_llm.layers.linear.Linear class method) (tensorrt_llm.layers.linear.LinearBase class method) (tensorrt_llm.layers.linear.RowLinear class method) trace_headers (tensorrt_llm.llmapi.RequestOutput attribute) transceiver_runtime (tensorrt_llm.llmapi.CacheTransceiverConfig attribute) transfer_mode (tensorrt_llm.llmapi.KvCacheRetentionConfig property) translate() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) transpose() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) trtllm-bench command line option --log_level --model --model_path --no-telemetry --revision --telemetry --workspace -m -w trtllm-bench-build command line option --dataset --max_batch_size --max_num_tokens --max_seq_len --no_weights_loading --pp_size --quantization --target_input_len --target_output_len --tp_size --trust_remote_code -pp -q -tp trtllm-bench-latency command line option --backend --beam_width --concurrency --config --custom_tokenizer --dataset --engine_dir --ep --extra_llm_api_options --iteration_log --kv_cache_free_gpu_mem_fraction --max_input_len --max_seq_len --medusa_choices --modality --num_requests --pp --report_json --sampler_options --tp --warmup trtllm-bench-throughput command line option --backend --beam_width --cluster_size --concurrency --config --custom_module_dirs --custom_tokenizer --data_device --dataset --disable_chunked_context --enable_chunked_context --engine_dir --eos_id --ep --extra_llm_api_options --image_data_format --iteration_log --kv_cache_free_gpu_mem_fraction --max_batch_size --max_input_len --max_num_tokens --max_seq_len --modality --no_skip_tokenizer_init --num_requests --output_json --pp --report_json --request_json --sampler_options --scheduler_policy --streaming --target_input_len --target_output_len --tp --warmup trtllm-eval command line option --backend --config --custom_tokenizer --disable_kv_cache_reuse --ep_size --extra_llm_api_options --gpus_per_node --kv_cache_free_gpu_memory_fraction --log_level --max_batch_size --max_beam_width --max_num_tokens --max_seq_len --model --no-telemetry --pp_size --revision --telemetry --tokenizer --tp_size --trust_remote_code trtllm-eval-aime25 command line option --apply_chat_template --chat_template_kwargs --dataset_path --fewshot_as_multiturn --log_samples --max_input_length --max_output_length --num_samples --output_dir --output_path --random_seed --sampling_seed --system_prompt --temperature --top_k --top_p trtllm-eval-aime26 command line option --apply_chat_template --chat_template_kwargs --dataset_path --fewshot_as_multiturn --log_samples --max_input_length --max_output_length --num_samples --output_dir --output_path --random_seed --sampling_seed --system_prompt --temperature --top_k --top_p trtllm-eval-cnn_dailymail command line option --apply_chat_template --dataset_path --max_input_length --max_output_length --num_samples --output_dir --random_seed --rouge_path --system_prompt trtllm-eval-covost2 command line option --apply_chat_template --dataset_path --dump_samples_path --lang_pair --max_input_length --max_output_length --num_samples --output_dir --random_seed --system_prompt --temperature trtllm-eval-gpqa_diamond command line option --apply_chat_template --chat_template_kwargs --dataset_path --log_samples --max_input_length --max_output_length --num_samples --output_dir --output_path --random_seed --sampling_seed --system_prompt --temperature --top_k --top_p trtllm-eval-gpqa_extended command line option --apply_chat_template --chat_template_kwargs --dataset_path --log_samples --max_input_length --max_output_length --num_samples --output_dir --output_path --random_seed --sampling_seed --system_prompt --temperature --top_k --top_p trtllm-eval-gpqa_main command line option --apply_chat_template --chat_template_kwargs --dataset_path --log_samples --max_input_length --max_output_length --num_samples --output_dir --output_path --random_seed --sampling_seed --system_prompt --temperature --top_k --top_p trtllm-eval-gsm8k command line option --apply_chat_template --chat_template_kwargs --dataset_path --fewshot_as_multiturn --log_samples --max_input_length --max_output_length --num_samples --output_dir --output_path --random_seed --sampling_seed --system_prompt --temperature --top_k --top_p trtllm-eval-json_mode_eval command line option --dataset_path --max_input_length --max_output_length --num_samples --output_dir --random_seed --system_prompt trtllm-eval-longbench_v1 command line option --apply_chat_template --chat_template_kwargs --dataset_path --log_samples --num_samples --output_dir --output_path --random_seed --system_prompt trtllm-eval-longbench_v2 command line option --apply_chat_template --chat_template_kwargs --cot --dataset_path --difficulty --domain --length --max_input_length --max_len --max_output_length --no_context --num_samples --output_dir --prompts_dir --rag --random_seed --start_idx --system_prompt --temperature --top_p trtllm-eval-mmlu command line option --accuracy_threshold --apply_chat_template --chat_template_kwargs --check_accuracy --dataset_path --max_input_length --max_output_length --num_fewshot --num_samples --output_dir --random_seed --system_prompt trtllm-eval-mmmu command line option --chat_template_kwargs --dataset_path --log_samples --max_input_length --max_output_length --num_samples --output_dir --output_path --random_seed --system_prompt trtllm-eval-mmmu_pro command line option --chat_template_kwargs --dataset_path --log_samples --max_input_length --max_output_length --num_samples --output_dir --output_path --random_seed --sampling_seed --subset --system_prompt --temperature --top_k --top_p trtllm-serve-disaggregated command line option --config --config_file --log_level --metadata_server_config_file --metrics-log-interval --request_timeout --schedule_style --server_start_timeout -c -l -m -r -s -t trtllm-serve-disaggregated_mpi_worker command line option --config --config_file --log_level -c trtllm-serve-mm_embedding_serve command line option --config --extra_encoder_options --free_gpu_memory_fraction --gpus_per_node --hf_revision --host --log_level --max_batch_size --max_num_tokens --metadata_server_config_file --no-telemetry --port --revision --telemetry --tensor_parallel_size --tp_size --trust_remote_code MODEL trtllm-serve-serve command line option --agent_percentage --agent_types --backend --chat_template --cluster_size --config --context_parallel_size --cp_size --custom_module_dirs --custom_tokenizer --disagg_cluster_uri --enable_attention_dp --enable_chunked_prefill --ep_size --extra_llm_api_options --extra_visual_gen_options --fail_fast_on_attention_window_too_large --free_gpu_memory_fraction --gpus_per_node --grpc --hf_revision --host --kv_cache_dtype --kv_cache_free_gpu_memory_fraction --log_level --max_batch_size --max_beam_width --max_num_tokens --max_seq_len --media_io_kwargs --metadata_server_config_file --middleware --moe_cluster_parallel_size --moe_expert_parallel_size --no-telemetry --num_postprocess_workers --otlp_traces_endpoint --pipeline_parallel_size --port --pp_size --reasoning_parser --revision --served_model_name --server_role --telemetry --tensor_parallel_size --tokenizer --tool_parser --tp_size --trust_remote_code --video_pruning_rate MODEL trtllm_modules_to_hf_modules (tensorrt_llm.runtime.ModelConfig attribute) TrtLlmArgs (class in tensorrt_llm.llmapi) TrtLlmArgs.Config (class in tensorrt_llm.llmapi) truncate_prompt_tokens (tensorrt_llm.llmapi.SamplingParams attribute) trust_remote_code (tensorrt_llm.llmapi.TorchLlmArgs attribute) (tensorrt_llm.llmapi.TrtLlmArgs attribute) TWOSHOT (tensorrt_llm.functional.AllReduceStrategy attribute) U UB (tensorrt_llm.functional.AllReduceStrategy attribute) unary() (in module tensorrt_llm.functional) unbind() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) unfuse_qkv_projections() (tensorrt_llm.models.SD3Transformer2DModel method) unpatchify() (tensorrt_llm.models.DiT method) unsqueeze() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) update() (tensorrt_llm.runtime.SamplingConfig method) update_forward_refs() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) update_kv_cache_type() (tensorrt_llm.llmapi.BuildConfig method) update_output_ids_by_offset() (tensorrt_llm.runtime.GenerationSession method) update_strategy() (tensorrt_llm.functional.AllReduceParams method) update_weights() (tensorrt_llm.llmapi.AsyncLLM method) upper() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method) use_beam_hyps (tensorrt_llm.runtime.SamplingConfig attribute) use_beam_search (tensorrt_llm.llmapi.SamplingParams attribute) use_cute_dsl_bf16_bmm (tensorrt_llm.llmapi.TorchLlmArgs attribute) use_cute_dsl_bf16_gemm (tensorrt_llm.llmapi.TorchLlmArgs attribute) use_cute_dsl_blockscaling_bmm (tensorrt_llm.llmapi.TorchLlmArgs attribute) use_cute_dsl_blockscaling_mm (tensorrt_llm.llmapi.TorchLlmArgs attribute) use_cute_dsl_paged_mqa_logits (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) use_cute_dsl_topk (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig attribute) use_dynamic_tree (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) use_gemm_allreduce_plugin (tensorrt_llm.runtime.GenerationSession property) use_gpt_attention_plugin (tensorrt_llm.runtime.GenerationSession property) use_kv_cache (tensorrt_llm.runtime.GenerationSession property) use_kv_cache_manager_v2 (tensorrt_llm.llmapi.KvCacheConfig attribute) use_lora() (tensorrt_llm.models.DecoderModel method) (tensorrt_llm.models.EncoderModel method) (tensorrt_llm.models.GemmaForCausalLM method) (tensorrt_llm.models.GPTForCausalLM method) (tensorrt_llm.models.LLaMAForCausalLM method) (tensorrt_llm.models.MLLaMAForCausalLM method) (tensorrt_llm.models.Phi3ForCausalLM method) (tensorrt_llm.models.PhiForCausalLM method) use_lora_plugin (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelRunner property) use_low_precision_moe_combine (tensorrt_llm.llmapi.MoeConfig attribute) use_mamba_conv1d_plugin (tensorrt_llm.runtime.GenerationSession property) use_meta_recipe (tensorrt_llm.llmapi.QuantConfig attribute) use_mrope (tensorrt_llm.llmapi.BuildConfig attribute) use_mtp_vanilla (tensorrt_llm.llmapi.MTPDecodingConfig attribute) use_prompt_tuning() (tensorrt_llm.models.EncoderModel method) use_python_scheduler (tensorrt_llm.llmapi.SchedulerConfig attribute) use_refit (tensorrt_llm.llmapi.BuildConfig attribute) use_rejection_sampling (tensorrt_llm.llmapi.AutoDecodingConfig attribute) (tensorrt_llm.llmapi.DFlashDecodingConfig attribute) (tensorrt_llm.llmapi.DraftTargetDecodingConfig attribute) (tensorrt_llm.llmapi.Eagle3DecodingConfig attribute) (tensorrt_llm.llmapi.EagleDecodingConfig attribute) (tensorrt_llm.llmapi.LookaheadDecodingConfig attribute) (tensorrt_llm.llmapi.MedusaDecodingConfig attribute) (tensorrt_llm.llmapi.MTPDecodingConfig attribute) (tensorrt_llm.llmapi.NGramDecodingConfig attribute) (tensorrt_llm.llmapi.PARDDecodingConfig attribute) (tensorrt_llm.llmapi.SADecodingConfig attribute) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) (tensorrt_llm.llmapi.UserProvidedDecodingConfig attribute) use_relaxed_acceptance_for_thinking (tensorrt_llm.llmapi.MTPDecodingConfig attribute) use_strip_plan (tensorrt_llm.llmapi.BuildConfig attribute) use_uvm (tensorrt_llm.llmapi.KvCacheConfig attribute) USER_PROVIDED (tensorrt_llm.models.SpeculativeDecodingMode attribute) UserProvidedDecodingConfig (class in tensorrt_llm.llmapi) UserProvidedDecodingConfig.Config (class in tensorrt_llm.llmapi) V validate() (tensorrt_llm.llmapi.AttentionDpConfig class method) (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.BuildCacheConfig class method) (tensorrt_llm.llmapi.BuildConfig class method) (tensorrt_llm.llmapi.CacheTransceiverConfig class method) (tensorrt_llm.llmapi.CalibConfig class method) (tensorrt_llm.llmapi.CudaGraphConfig class method) (tensorrt_llm.llmapi.DeepSeekSparseAttentionConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.DynamicBatchConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.ExtendedRuntimePerfKnobConfig class method) (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MoeConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) (tensorrt_llm.llmapi.QuantConfig class method) (tensorrt_llm.llmapi.ReorderRequestPolicyConfig class method) (tensorrt_llm.llmapi.RocketSparseAttentionConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SAEnhancerConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.SchedulerConfig class method) (tensorrt_llm.llmapi.SkipSoftmaxAttentionConfig class method) (tensorrt_llm.llmapi.TorchCompileConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) validate_and_init_tokenizer() (tensorrt_llm.llmapi.TorchLlmArgs method) (tensorrt_llm.llmapi.TrtLlmArgs method) validate_attention_dp_config() (tensorrt_llm.llmapi.AttentionDpConfig method) validate_build_config_remaining() (tensorrt_llm.llmapi.TrtLlmArgs method) validate_build_config_with_runtime_params() (tensorrt_llm.llmapi.TrtLlmArgs method) validate_capture_num_tokens() (tensorrt_llm.llmapi.TorchCompileConfig class method) validate_checkpoint_format() (tensorrt_llm.llmapi.TorchLlmArgs method) validate_cuda_graph_config() (tensorrt_llm.llmapi.CudaGraphConfig method) validate_cute_dsl_bf16() (tensorrt_llm.llmapi.TorchLlmArgs method) validate_draft_len_schedule_and_sort() (tensorrt_llm.llmapi.AutoDecodingConfig class method) (tensorrt_llm.llmapi.DFlashDecodingConfig class method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig class method) (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) (tensorrt_llm.llmapi.LookaheadDecodingConfig class method) (tensorrt_llm.llmapi.MedusaDecodingConfig class method) (tensorrt_llm.llmapi.MTPDecodingConfig class method) (tensorrt_llm.llmapi.NGramDecodingConfig class method) (tensorrt_llm.llmapi.PARDDecodingConfig class method) (tensorrt_llm.llmapi.SADecodingConfig class method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig class method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig class method) validate_draft_target_config() (tensorrt_llm.llmapi.DraftTargetDecodingConfig method) validate_dtype() (tensorrt_llm.llmapi.KvCacheConfig class method) (tensorrt_llm.llmapi.TorchLlmArgs class method) (tensorrt_llm.llmapi.TrtLlmArgs class method) validate_eagle_choices() (tensorrt_llm.llmapi.Eagle3DecodingConfig class method) (tensorrt_llm.llmapi.EagleDecodingConfig class method) validate_eagle_config() (tensorrt_llm.llmapi.Eagle3DecodingConfig method) (tensorrt_llm.llmapi.EagleDecodingConfig method) validate_early_first_token_response() (tensorrt_llm.llmapi.TorchLlmArgs method) validate_enable_build_cache() (tensorrt_llm.llmapi.TrtLlmArgs method) validate_free_gpu_memory_fraction() (tensorrt_llm.llmapi.KvCacheConfig class method) validate_gpus_per_node() (tensorrt_llm.llmapi.TorchLlmArgs class method) (tensorrt_llm.llmapi.TrtLlmArgs class method) validate_helix_tokens_per_block() (tensorrt_llm.llmapi.TorchLlmArgs method) validate_histogram_buckets() (tensorrt_llm.llmapi.PrometheusMetricsConfig class method) validate_kv_cache_dtype() (tensorrt_llm.llmapi.TrtLlmArgs method) validate_load_balancer() (tensorrt_llm.llmapi.TorchLlmArgs method) validate_lora_config_consistency() (tensorrt_llm.llmapi.TorchLlmArgs method) (tensorrt_llm.llmapi.TrtLlmArgs method) validate_max_attention_window() (tensorrt_llm.llmapi.KvCacheConfig class method) validate_max_concurrency_and_draft_len_schedule_mutually_exclusive() (tensorrt_llm.llmapi.AutoDecodingConfig method) (tensorrt_llm.llmapi.DFlashDecodingConfig method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig method) (tensorrt_llm.llmapi.Eagle3DecodingConfig method) (tensorrt_llm.llmapi.EagleDecodingConfig method) (tensorrt_llm.llmapi.LookaheadDecodingConfig method) (tensorrt_llm.llmapi.MedusaDecodingConfig method) (tensorrt_llm.llmapi.MTPDecodingConfig method) (tensorrt_llm.llmapi.NGramDecodingConfig method) (tensorrt_llm.llmapi.PARDDecodingConfig method) (tensorrt_llm.llmapi.SADecodingConfig method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig method) validate_max_gpu_total_bytes() (tensorrt_llm.llmapi.KvCacheConfig class method) validate_max_util_for_resume() (tensorrt_llm.llmapi.KvCacheConfig class method) validate_model_format_misc() (tensorrt_llm.llmapi.TrtLlmArgs method) validate_mx_config() (tensorrt_llm.llmapi.TorchLlmArgs method) validate_ngram_config() (tensorrt_llm.llmapi.NGramDecodingConfig method) validate_parallel_config() (tensorrt_llm.llmapi.TorchLlmArgs method) (tensorrt_llm.llmapi.TrtLlmArgs method) validate_peft_cache_config() (tensorrt_llm.llmapi.TorchLlmArgs method) (tensorrt_llm.llmapi.TrtLlmArgs method) validate_ray_placement_config() (tensorrt_llm.llmapi.TorchLlmArgs method) validate_ray_worker_extension_cls() (tensorrt_llm.llmapi.TorchLlmArgs method) validate_rejection_sampling_config() (tensorrt_llm.llmapi.AutoDecodingConfig method) (tensorrt_llm.llmapi.DFlashDecodingConfig method) (tensorrt_llm.llmapi.DraftTargetDecodingConfig method) (tensorrt_llm.llmapi.Eagle3DecodingConfig method) (tensorrt_llm.llmapi.EagleDecodingConfig method) (tensorrt_llm.llmapi.LookaheadDecodingConfig method) (tensorrt_llm.llmapi.MedusaDecodingConfig method) (tensorrt_llm.llmapi.MTPDecodingConfig method) (tensorrt_llm.llmapi.NGramDecodingConfig method) (tensorrt_llm.llmapi.PARDDecodingConfig method) (tensorrt_llm.llmapi.SADecodingConfig method) (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig method) (tensorrt_llm.llmapi.UserProvidedDecodingConfig method) validate_runtime_args() (tensorrt_llm.llmapi.TorchLlmArgs method) (tensorrt_llm.llmapi.TrtLlmArgs method) validate_sa_config() (tensorrt_llm.llmapi.SADecodingConfig method) validate_speculative_beam_history_d2h() (tensorrt_llm.llmapi.TorchLlmArgs method) validate_speculative_config() (tensorrt_llm.llmapi.TorchLlmArgs method) (tensorrt_llm.llmapi.TrtLlmArgs method) validate_speculative_model() (tensorrt_llm.llmapi.Eagle3DecodingConfig method) (tensorrt_llm.llmapi.EagleDecodingConfig method) VERBATIM (tensorrt_llm.models.GemmaConfig attribute) video_preprocess() (tensorrt_llm.runtime.MultimodalModelRunner method) video_pruning_rate (tensorrt_llm.llmapi.TorchLlmArgs attribute) view() (in module tensorrt_llm.functional) (tensorrt_llm.functional.Tensor method) (tensorrt_llm.runtime.TensorInfo method) visual_engine_dir (tensorrt_llm.runtime.MultimodalModelRunner property) visualize_network (tensorrt_llm.llmapi.BuildConfig attribute) vocab_size (tensorrt_llm.runtime.GenerationSession property) (tensorrt_llm.runtime.ModelConfig attribute) (tensorrt_llm.runtime.ModelRunner property) (tensorrt_llm.runtime.ModelRunnerCpp property) vocab_size_padded (tensorrt_llm.runtime.ModelRunner property) (tensorrt_llm.runtime.ModelRunnerCpp property) W W4A16 (tensorrt_llm.llmapi.QuantAlgo attribute) W4A16_AWQ (tensorrt_llm.llmapi.QuantAlgo attribute) W4A16_GPTQ (tensorrt_llm.llmapi.QuantAlgo attribute) W4A16_MXFP4 (tensorrt_llm.llmapi.QuantAlgo attribute) W4A8_AWQ (tensorrt_llm.llmapi.QuantAlgo attribute) W4A8_MXFP4_FP8 (tensorrt_llm.llmapi.QuantAlgo attribute) W4A8_MXFP4_MXFP8 (tensorrt_llm.llmapi.QuantAlgo attribute) W4A8_NVFP4_FP8 (tensorrt_llm.llmapi.QuantAlgo attribute) W4A8_QSERVE_PER_CHANNEL (tensorrt_llm.llmapi.QuantAlgo attribute) W4A8_QSERVE_PER_GROUP (tensorrt_llm.llmapi.QuantAlgo attribute) W8A16 (tensorrt_llm.llmapi.QuantAlgo attribute) W8A16_GPTQ (tensorrt_llm.llmapi.QuantAlgo attribute) W8A8_SQ_PER_CHANNEL (tensorrt_llm.llmapi.QuantAlgo attribute) W8A8_SQ_PER_CHANNEL_PER_TENSOR_PLUGIN (tensorrt_llm.llmapi.QuantAlgo attribute) W8A8_SQ_PER_CHANNEL_PER_TOKEN_PLUGIN (tensorrt_llm.llmapi.QuantAlgo attribute) W8A8_SQ_PER_TENSOR_PER_TOKEN_PLUGIN (tensorrt_llm.llmapi.QuantAlgo attribute) W8A8_SQ_PER_TENSOR_PLUGIN (tensorrt_llm.llmapi.QuantAlgo attribute) waiting_queue_policy (tensorrt_llm.llmapi.SchedulerConfig attribute) warn_on_unstable_feature_usage() (tensorrt_llm.llmapi.TorchLlmArgs method) weight_loader() (tensorrt_llm.layers.attention.DeepseekV2Attention method) (tensorrt_llm.layers.embedding.Embedding method) (tensorrt_llm.layers.linear.LinearBase method) weight_sparsity (tensorrt_llm.llmapi.BuildConfig attribute) weight_streaming (tensorrt_llm.llmapi.BuildConfig attribute) where() (in module tensorrt_llm.functional) WhisperEncoder (class in tensorrt_llm.models) window_size (tensorrt_llm.llmapi.RocketSparseAttentionConfig attribute) with_traceback() (tensorrt_llm.llmapi.RequestError method) workspace (tensorrt_llm.llmapi.TrtLlmArgs attribute) wrapped_property (tensorrt_llm.llmapi.TorchLlmArgs attribute), [1] (tensorrt_llm.llmapi.TrtLlmArgs attribute) write_interval (tensorrt_llm.llmapi.SaveHiddenStatesDecodingConfig attribute) Y yarn (tensorrt_llm.functional.PositionEmbeddingType attribute) (tensorrt_llm.functional.RotaryScalingType attribute) Z zfill() (tensorrt_llm.llmapi.BatchingType method) (tensorrt_llm.llmapi.CapacitySchedulerPolicy method) (tensorrt_llm.llmapi.ContextChunkingPolicy method) (tensorrt_llm.llmapi.QuantAlgo method)