transcribe

Fixing S3 Bottlenecks: Scalable I/O for Ray with Alluxio | Ray Summit 2025

Anyscale · 9m · transcribed 14d ago
More from Anyscale Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Introduction to Aluxio

How can Aluxio help eliminate S3 bottlenecks?

Aluxio serves as a data acceleration layer that allows users to access data from various storage systems as if it were on a local disk, thus simplifying data handling for AI workloads.

  • Aluxio addresses the challenge of slow data access in AI applications.
  • It provides a plug-and-play solution for data access across different storage systems.
  • Users can focus on improving their models without worrying about data management issues.
# 1:54

Features of Aluxio

What are the key features of Aluxio?

Aluxio is a distributed system that allows easy mounting of multiple storage buckets, providing a unified caching and data acceleration service without requiring code changes.

  • Aluxio supports distributed systems with scalability from a few to thousands of workers.
  • It offers instant data availability and high throughput without data migration.
  • Users can access data seamlessly across different frameworks and storage types.
# 3:48

Performance Metrics

How does Aluxio improve data access performance?

Aluxio significantly enhances data access speeds and reduces latency compared to direct cloud storage access, achieving high throughput and low latency for workloads.

  • Aluxio can achieve up to 60 GB/s throughput with multiple workers, outperforming direct cloud storage access.
  • It provides consistent sub-millisecond latency for low-latency workloads.
  • Scalability is proven with large deployments managing petabytes of data.
# 5:43

Integration with Ray

How does Aluxio integrate with Ray?

Aluxio integrates with Ray to enhance data throughput, allowing users to access data easily without data migration, thus improving performance for AI applications.

  • Integration with Ray shows a 3.5x improvement in throughput.
  • Users can mount storage buckets quickly, enabling immediate access to data.
  • Aluxio supports various APIs for seamless integration with existing frameworks.
# 7:37

Use Cases and Collaborations

What are some practical applications of Aluxio?

Aluxio is being used in various applications, including integrating with Hugging Face for model loading and optimizing data access patterns for improved performance.

  • Aluxio enables fast model loading for AI applications, achieving terabytes per second throughput.
  • Collaborations with companies like Salesforce demonstrate Aluxio's capability in handling large datasets efficiently.
  • The service is designed to meet the demands of low-latency applications.

Transcript

0:03 Hi everyone, my name is Bin from Alux. Today I'm going to talk about how Alux can help you to eliminate S3 bottlenecks to scale your IO and how you can use this to potentially benefit your rear applications. Okay, so the problem statement ray is fast and you have fast GPUs. However, the data may not be fast. Okay. based on my experience working with the data modelers, AI researchers from time to time they only want to focus their own jobs in making the model better or making the inferencing better. They do not want to handle issues like oh I have to handle the data sets or it might be too large to fit my local disk or I'm downloading slow or unstable from the resource. had handled different URLs in different environment when I'm launching my jobs in different clusters or different clouds. It's handling network issues, permission issues, syncing back data from your remote storage. Okay, none of them like handle issues like this. They all prefer a easy, fast, scalable solution so that they can just start to use the data in the plugandplay manner. Okay, so this is where Alux comes. We are a data acceleration layer for array and for the AI workloads. You can think you have the persistent layer on S3, Oracle, Azure, GCS, even the on-remise storage like Hadoop and then you just put this data access and acceleration layer like Aluxio put into your GPU clusters or CPU clusters and then you can access data just like accessing your local disk, your local mount. gives you the feeling you are just using a NAS but actually it's backed by a storage or a cloud storage service a buckets okay that is in a very high level what Alexio is just give you a feeling this is but basically aluxio the sequence sequence if you want to use this it's transparent and easy admin can just mount a buckets or multiple different buckets into aluxio and then I'm showing the pics way okay then You can start to use your data which is backed in the persistent buckets and just look like they are sitting in your local disk and you can read data, you can write data and remember this is a actually yeah just recall this is actually not a single node solution. We are talking about distributed system. So aluxo is not like a S3 FS or GCS FS if you have heard this tools before. This is a distributed system that runs can be a few workers to thousands of workers. So that presents the unified and shared caching and data acceleration service for you. So what means to you as a AI user or customer get a transparent simple regardless data cloud or storage you're using no code change no configuration change works with most of these existing frameworks and the data will be available instantly once they are pushed into the buckets no wait for data to be copied to your local storage or to somewhere as I mentioned this is a distributed system so you actually get the aggregated performance cluster wise if you have hundreds of Alexa workers you get terabytes per second as aggregated throughput and the time to first bite latency can be sums I will show you the benchmark more importantly this is proven the scalability scalability is proven in production we have a large customers running pabytes multiple pabytes in the cache volume with billions of files and beyond cached okay so this is whenever the benchmark you are reading the data with six from six aluxy workers where data sits in a cloud object storage. If you go to object storage directly you get around like for example in this case 7 gigabytes per second to pull the data directly from the object store and this is about the same or slight less than if you have a one Alexa single worker which gives you 10 gigabytes per second. Note that this is already close to 80 more than 80% of the utilization of the nick which is 100 gigabits nick but you have six alloxy workers putting together you get a more than like close to 60 gig bytes per second and in terms of IOPS per second is similar like we're able to get a 5x IOPS per second compared to one worker when there six workers but and this is like almost 20x more than if you just talking to a cloud storage directly.

4:55 Latency, my favorite part. Basically, Aluxio is designed for low latency workloads. If you have really low latency, latency sensitive workloads, you want to get the data from your cloud storage fast. as you can see, if you directly talk to the cloud storage, it gives you on average about 20 milliseconds. This is average. If you're talking about P99, this can be three or 200 milliseconds. But with Aluxo it gives you consistent stable submillconds latency for both average and P99 proven scalability. you can run this with one terabytes per second with individual cluster with enough number of nodes and the largest single cluster manages more than three pabytes of data and a single cluster we manages more than billion files. Why this matters?

5:45 when I'm talking to people are training their models for self-driving cars or robotics, they're often using mill tens of millions, hundred millions or even billions of files, small images or video clips in their data sets. Okay, 50 production clusters deployed globally for the largest customer. Okay, so for Ray, I'm talking about Ray here today. we have an integration with ray and you can use our fs back interface to using ray to talk to aluxio instead of talking to like for example ray data and we shows there's a 3.5x improvement in terms of throughput from the same region okay so the transparent and easy access will enable this integration for ray plus or luxio as you can see you can have different type of frameworks here and you and choose your favorite APIs, S3 API, PZIX like or Python for R. We support like we basically can use any of these APIs to talk to this caching service which is backed by the storage and the user experience will be like hey I just mount this bucket my bucket mounted into the aluxo directory and then the metadata this is only a metadata operation there's no data migration and this can finish in a second or two and then you can have two different ways your re applications can now just read from a local directory this is you can view this as a PVC persistent volume claim to your applications and you can start to read and write from the same persistent volume and we also provide as I mentioned like you can use S3 protocol. So you can set the endpoints to be aluxy S3 endpoints instead instead of the default for example S3 endpoints or some other object store endpoints then you can still continue to use your S3 APIs to engage with Alexio.

7:42 Yeah, I want to share something fun like we're doing recently. one is we are integrating with hugging face. So if you have a re applications or some other things you can just use and Luxio and mount hugging face hub model hub as a file system and after this like for example running the first command you can directly go to get these models just like they are sitting on your Linux loader folder locally and after that like we also have certain optimizations to for you to load safe tensor formats from hugging face into GPUs and the reason behind the philosophy behind is we are doing some anal we did some analysis on the data access pattern to this safe tensor formats.

8:26 So once you touch part of this it's it's a format it has many attri quite a few files in the same directory for the safe tensor files and then we can predict like what is the next one or next next one for you to access and this gives you pretty good rates to load data into your GPUs. Okay. So like for example we enable fireworks to load models fast to avoid this code start problems in seconds and the we get like we see terabytes per second aggregated throughput in a single deployment for this model inferencing use case and for for example Salesforce they are building a sumoc latency for the agentic memory layer where they having this pabytes of data sitting on S3 but they want really sumosconds So we are able to help them.

9:16 We have a joint white paper talking about what we do there. that concludes my talk. So we have a booths and if you want labubu you can feel free to scan this and come to our booths. I'm happy to talk to you.

Summary

Aluxio provides a data acceleration layer designed to eliminate bottlenecks in S3 and other cloud storage systems, enabling faster data access for AI and machine learning workloads. By integrating seamlessly with existing frameworks, Aluxio allows users to access large datasets as if they were local files, significantly improving data throughput and reducing latency.

- Aluxio acts as a data acceleration layer for various storage systems, including S3, Oracle, and Azure.
- It enables users to access data in a plug-and-play manner without needing to manage complex data handling issues.
- The system supports distributed architecture, allowing for high throughput and low latency across multiple workers.
- Users can achieve terabytes per second in aggregated throughput and sub-millisecond latency for data access.
- Aluxio integrates with frameworks like Ray, improving throughput by 3.5 times compared to direct cloud storage access.
- The platform supports various APIs, including S3 and Python, making it versatile for different applications.
- Recent integrations include Hugging Face, allowing users to access models as if they were local files, enhancing data loading speeds.
- Aluxio has proven scalability, managing petabytes of data and billions of files across multiple production clusters.

Questions Answered

How can Aluxio help eliminate S3 bottlenecks?

Aluxio serves as a data acceleration layer that allows users to access data from various storage systems as if it were on a local disk, thus simplifying data handling for AI workloads.

What are the key features of Aluxio?

Aluxio is a distributed system that allows easy mounting of multiple storage buckets, providing a unified caching and data acceleration service without requiring code changes.

How does Aluxio improve data access performance?

Aluxio significantly enhances data access speeds and reduces latency compared to direct cloud storage access, achieving high throughput and low latency for workloads.

How does Aluxio integrate with Ray?

Aluxio integrates with Ray to enhance data throughput, allowing users to access data easily without data migration, thus improving performance for AI applications.

What are some practical applications of Aluxio?

Aluxio is being used in various applications, including integrating with Hugging Face for model loading and optimizing data access patterns for improved performance.

© transcribe · For agents Built with care and craft by Gokul Rajaram