AWS SageMaker accelerates large language model scaling
Amazon Web Services has introduced Fast Model Loader for Amazon SageMaker Inference. The new capability reduces model loading times and speeds up autoscaling for large language models (LLMs).

Amazon Web Services (AWS) has introduced Fast Model Loader for its Amazon SageMaker Inference service. The new feature is designed to significantly reduce the time required to load and scale large language models (LLMs) for inference.
As LLMs grow to hundreds of billions of parameters and require substantial memory, the process of loading these models onto accelerators has become a bottleneck. This challenge has hindered efficient deployment and rapid scaling of LLMs to handle fluctuating traffic patterns.
The Fast Model Loader reportedly works by streaming model weights directly from Amazon Simple Storage Service (S3) to the accelerator, reducing load times. AWS states this can lead to up to a 19% reduction in latency when scaling a new model copy onto a new instance for inference.
This new capability is part of the SageMaker Large Model Inference (LMI) offering, aimed at helping customers deploy LLMs on SageMaker Inference more quickly. AWS announced the technology at its AWS re:Invent 2024 event.