Close Menu
StreamLineCrypto.comStreamLineCrypto.com
  • Home
  • Crypto News
  • Bitcoin
  • Altcoins
  • NFT
  • Defi
  • Blockchain
  • Metaverse
  • Regulations
  • Trading
What's Hot

Strategy keeps STRC dividend at 12% below $90

August 2, 2026

A massive stablecoin fragmentation war is brewing between tech giants and a startup is aiming to capitalize on it

August 2, 2026

The US just blacklisted the Iranian maritime scheme forcing commercial ships to pay Bitcoin tolls for safe passage

August 2, 2026
Facebook X (Twitter) Instagram
Sunday, August 2 2026
  • Contact Us
  • Privacy Policy
  • Cookie Privacy Policy
  • Terms of Use
  • DMCA
Facebook X (Twitter) Instagram
StreamLineCrypto.comStreamLineCrypto.com
  • Home
  • Crypto News
  • Bitcoin
  • Altcoins
  • NFT
  • Defi
  • Blockchain
  • Metaverse
  • Regulations
  • Trading
StreamLineCrypto.comStreamLineCrypto.com

NVIDIA Enhances Llama 3.1 405B Performance with TensorRT Model Optimizer

August 29, 2024Updated:August 29, 2024No Comments4 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr Email
NVIDIA Enhances Llama 3.1 405B Performance with TensorRT Model Optimizer
Share
Facebook Twitter LinkedIn Pinterest Email
ad


Lawrence Jengar
Aug 29, 2024 16:10

NVIDIA’s TensorRT Mannequin Optimizer considerably boosts efficiency of Meta’s Llama 3.1 405B massive language mannequin on H200 GPUs.





Meta’s Llama 3.1 405B massive language mannequin (LLM) is attaining new ranges of efficiency because of NVIDIA’s TensorRT Mannequin Optimizer, in response to the NVIDIA Technical Weblog. The enhancements have resulted in as much as a 1.44x improve in throughput when working on NVIDIA H200 GPUs.

Excellent Llama 3.1 405B Inference Throughput with TensorRT-LLM

TensorRT-LLM has already delivered exceptional inference throughput for Llama 3.1 405B because the mannequin’s launch. This was achieved by way of varied optimizations, together with in-flight batching, KV caching, and optimized consideration kernels. These strategies have accelerated inference efficiency whereas sustaining decrease precision compute.

TensorRT-LLM added assist for the official Llama FP8 quantization recipe, which calculates static and dynamic scaling elements to protect most accuracy. Moreover, user-defined kernels reminiscent of matrix multiplications from FBGEMM are optimized by way of plug-ins inserted into the community graph at compile time.

Boosting Efficiency As much as 1.44x with TensorRT Mannequin Optimizer

NVIDIA’s customized FP8 post-training quantization (PTQ) recipe, accessible by way of the TensorRT Mannequin Optimizer library, enhances Llama 3.1 405B throughput and reduces latency with out sacrificing accuracy. This recipe incorporates FP8 KV cache quantization and self-attention static quantization, lowering inference compute overhead.

Desk 1 demonstrates the utmost throughput efficiency, exhibiting vital enhancements throughout varied enter and output sequence lengths on an 8-GPU HGX H200 system. The system options eight NVIDIA H200 Tensor Core GPUs with 141 GB of HBM3e reminiscence every and 4 NVLink Switches, offering 900 GB/s of GPU-to-GPU bandwidth.








Most Throughput Efficiency – Output Tokens/Second
8 NVIDIA H200 Tensor Core GPUs
Enter | Output Sequence Lengths2,048 | 12832,768 | 2,048120,000 | 2,048
TensorRT Mannequin Optimizer FP8463.1320.171.5
Official Llama FP8 Recipe399.9230.849.6
Speedup1.16x1.39x1.44x

Desk 1. Most throughput efficiency of Llama 3.1 405B with NVIDIA inside measurements

Equally, Desk 2 presents the minimal latency efficiency utilizing the identical enter and output sequence lengths.








Batch Measurement = 1 Efficiency – Output Tokens/Second
8 NVIDIA H200 Tensor Core GPUs
Enter | Output Sequence Lengths2,048 | 12832,768 | 2,048120,000 | 2,048
TensorRT Mannequin Optimizer FP849.644.227.2
Official Llama FP8 Recipe37.433.122.8
Speedup1.33x1.33x1.19x

Desk 2. Minimal latency efficiency of Llama 3.1 405B with NVIDIA inside measurements

These outcomes point out that H200 GPUs with TensorRT-LLM and TensorRT Mannequin Optimizer are delivering superior efficiency in each latency-optimized and throughput-optimized situations. The TensorRT Mannequin Optimizer FP8 recipe additionally achieved comparable accuracy with the official Llama 3.1 FP8 recipe on the Massively Multitask Language Understanding (MMLU) and MT-Bench benchmarks.

Becoming Llama 3.1 405B on Simply Two H200 GPUs with INT4 AWQ

For builders with {hardware} useful resource constraints, the INT4 AWQ approach in TensorRT Mannequin Optimizer compresses the mannequin, permitting Llama 3.1 405B to suit on simply two H200 GPUs. This methodology reduces the required reminiscence footprint considerably by compressing the weights all the way down to 4-bit integers whereas encoding activations utilizing FP16.

Tables 4 and 5 present the utmost throughput and minimal latency efficiency measurements, demonstrating that the INT4 AWQ methodology gives comparable accuracy scores to the Llama 3.1 official FP8 recipe from Meta.






Most Throughput Efficiency – Output Tokens/Second
2 NVIDIA H200 Tensor Core GPUs
Enter | Output Sequence Lengths2,048 | 12832,768 | 2,04860,000 | 2,048
TensorRT Mannequin Optimizer INT4 AWQ75.628.716.2

Desk 4. Most throughput efficiency of Llama 3.1 405B with NVIDIA inside measurements






Batch Measurement = 1 Efficiency – Output Tokens/Second
2 NVIDIA H200 Tensor Core GPUs
Enter | Output Sequence Lengths2,048 | 12832,768 | 2,04860,000 | 2,048
TensorRT Mannequin Optimizer INT4 AWQ21.618.712.8

Desk 5. Minimal latency efficiency of Llama 3.1 405B with NVIDIA inside measurements

NVIDIA’s developments in TensorRT Mannequin Optimizer and TensorRT-LLM are paving the way in which for enhanced efficiency and effectivity in working massive language fashions like Llama 3.1 405B. These enhancements supply builders extra flexibility and cost-efficiency, whether or not they have intensive {hardware} sources or extra constrained environments.

Picture supply: Shutterstock


ad
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Related Posts

A massive stablecoin fragmentation war is brewing between tech giants and a startup is aiming to capitalize on it

August 2, 2026

The US just blacklisted the Iranian maritime scheme forcing commercial ships to pay Bitcoin tolls for safe passage

August 2, 2026

Strategy Holds Preferred STRC Dividend at 12% as Price Still Below Par

August 2, 2026

New York asks judge to force Kalshi to hand over the names, wagers, and losses of its local bettors

August 2, 2026
Add A Comment
Leave A Reply Cancel Reply

ad
What's New Here!
Strategy keeps STRC dividend at 12% below $90
August 2, 2026
A massive stablecoin fragmentation war is brewing between tech giants and a startup is aiming to capitalize on it
August 2, 2026
The US just blacklisted the Iranian maritime scheme forcing commercial ships to pay Bitcoin tolls for safe passage
August 2, 2026
Strategy Holds Preferred STRC Dividend at 12% as Price Still Below Par
August 2, 2026
Minnesota loses first round against Kalshi, Polymarket
August 2, 2026
Facebook X (Twitter) Instagram Pinterest
  • Contact Us
  • Privacy Policy
  • Cookie Privacy Policy
  • Terms of Use
  • DMCA
© 2026 StreamlineCrypto.com - All Rights Reserved!

Type above and press Enter to search. Press Esc to cancel.