Close Menu
StreamLineCrypto.comStreamLineCrypto.com
  • Home
  • Crypto News
  • Bitcoin
  • Altcoins
  • NFT
  • Defi
  • Blockchain
  • Metaverse
  • Regulations
  • Trading
What's Hot

Ethereum Proposal to Slash Staking Rewards Sparks Backlash

August 5, 2026

GameStop plans $1.4 billion stock swap as Bitcoin collateral risk emerges

August 4, 2026

US, UK deepen stablecoin talks after GENIUS Act

August 4, 2026
Facebook X (Twitter) Instagram
Wednesday, August 5 2026
  • Contact Us
  • Privacy Policy
  • Cookie Privacy Policy
  • Terms of Use
  • DMCA
Facebook X (Twitter) Instagram
StreamLineCrypto.comStreamLineCrypto.com
  • Home
  • Crypto News
  • Bitcoin
  • Altcoins
  • NFT
  • Defi
  • Blockchain
  • Metaverse
  • Regulations
  • Trading
StreamLineCrypto.comStreamLineCrypto.com

IBM Research Unveils Innovations to Accelerate Enterprise AI Training

September 23, 2024Updated:September 23, 2024No Comments3 Mins Read
Facebook Twitter Pinterest LinkedIn Tumblr Email
IBM Research Unveils Innovations to Accelerate Enterprise AI Training
Share
Facebook Twitter LinkedIn Pinterest Email
ad


Zach Anderson
Sep 23, 2024 03:32

IBM Analysis introduces new knowledge processing methods to expedite AI mannequin coaching utilizing CPU assets, considerably enhancing effectivity.





IBM Analysis has unveiled groundbreaking improvements aimed toward scaling the information processing pipeline for enterprise AI coaching, in response to IBM Analysis. These developments are designed to expedite the creation of highly effective AI fashions, reminiscent of IBM’s Granite fashions, by leveraging the plentiful capability of CPUs.

Optimizing Information Preparation

Earlier than coaching AI fashions, huge quantities of knowledge have to be ready. This knowledge usually comes from various sources like web sites, PDFs, and information articles, and should bear a number of preprocessing steps. These steps embody filtering out irrelevant HTML code, eradicating duplicates, and screening for abusive content material. These duties, although important, are usually not constrained by the provision of GPUs.

Petros Zerfos, IBM Analysis’s principal analysis scientist for watsonx knowledge engineering, emphasised the significance of environment friendly knowledge processing. “A big a part of the effort and time that goes into coaching these fashions is making ready the information for these fashions,” Zerfos mentioned. His staff has been creating strategies to reinforce the effectivity of knowledge processing pipelines, drawing experience from numerous domains together with pure language processing, distributed computing, and storage techniques.

Leveraging CPU Capability

Many steps within the knowledge processing pipeline contain “embarrassingly parallel” computations, permitting every doc to be processed independently. This parallel processing can considerably pace up knowledge preparation by distributing duties throughout quite a few CPUs. Nonetheless, some steps, reminiscent of eradicating duplicate paperwork, require entry to the whole dataset, which can’t be carried out in parallel.

To speed up IBM’s Granite mannequin improvement, the staff has developed processes to quickly provision and make the most of tens of 1000’s of CPUs. This method includes marshalling idle CPU capability throughout IBM’s Cloud datacenter community, guaranteeing excessive communication bandwidth between CPUs and knowledge storage. Conventional object storage techniques usually trigger CPUs to idle as a result of low efficiency; thus, the staff employed IBM’s high-performance Storage Scale file system to cache lively knowledge effectively.

Scaling Up AI Coaching

Over the previous yr, IBM has scaled as much as 100,000 vCPUs within the IBM Cloud, processing 14 petabytes of uncooked knowledge to provide 40 trillion tokens for AI mannequin coaching. The staff has automated these knowledge pipelines utilizing Kubeflow on IBM Cloud. Their strategies have confirmed to be 24 instances quicker in processing knowledge from Widespread Crawl in comparison with earlier methods.

All of IBM’s open-sourced Granite code and language fashions have been educated utilizing knowledge ready by means of these optimized pipelines. Moreover, IBM has made vital contributions to the AI group by creating the Information Prep Equipment, a toolkit hosted on GitHub. This equipment streamlines knowledge preparation for giant language mannequin purposes, supporting pre-training, fine-tuning, and retrieval-augmented technology (RAG) use circumstances. Constructed on distributed processing frameworks like Spark and Ray, the equipment permits builders to construct scalable customized modules.

For extra info, go to the official IBM Analysis weblog.

Picture supply: Shutterstock


ad
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Related Posts

Ethereum Proposal to Slash Staking Rewards Sparks Backlash

August 5, 2026

GameStop plans $1.4 billion stock swap as Bitcoin collateral risk emerges

August 4, 2026

Self Custody Is Dead. Long Live Self Custody

August 4, 2026

Hester ‘Crypto Mom’ Peirce Optimistic About Clarity Act

August 4, 2026
Add A Comment
Leave A Reply Cancel Reply

ad
What's New Here!
Ethereum Proposal to Slash Staking Rewards Sparks Backlash
August 5, 2026
GameStop plans $1.4 billion stock swap as Bitcoin collateral risk emerges
August 4, 2026
US, UK deepen stablecoin talks after GENIUS Act
August 4, 2026
SpaceX taps NVIDIA for 1M-satellite AI plan
August 4, 2026
Self Custody Is Dead. Long Live Self Custody
August 4, 2026
Facebook X (Twitter) Instagram Pinterest
  • Contact Us
  • Privacy Policy
  • Cookie Privacy Policy
  • Terms of Use
  • DMCA
© 2026 StreamlineCrypto.com - All Rights Reserved!

Type above and press Enter to search. Press Esc to cancel.