‏إظهار الرسائل ذات التسميات Llama. إظهار كافة الرسائل
‏إظهار الرسائل ذات التسميات Llama. إظهار كافة الرسائل

Meta & Cerebras Unleash AI Speed—18x Faster Than GPU-based Solutions

Meta & Cerebras Unleash AI Speed—18x Faster Than GPU-based Solutions

Meta has officially teamed up with Cerebras Systems to supercharge its Llama API, delivering inference speeds up to 18 times faster than traditional GPU-based solutions. This move positions Meta to compete directly with OpenAI, Anthropic, and Google in the AI inference market, where developers purchase tokens to power their applications.

Cerebras Systems is a cutting-edge Al hardware company specializing in wafer-scale computing, designed to accelerate deep learning and Al inference. Their Wafer-Scale Engine (WSE) is the largest semiconductor chip ever built, offering unprecedented speed and efficiency compared to traditional GPUs.

The Cerebras system enables over 2,600 tokens per second for Llama 4 Scout, compared to 130 tokens per second for ChatGPT and 25 tokens per second for DeepSeek. This speed boost unlocks real-time AI applications, including low-latency voice systems, interactive code generation, and instant multi-step reasoning.

This collaboration positions Cerebras as a major player in Al infrastructure, challenging Nvidia's dominance in Al hardware.

Meta’s shift from just providing open-source models to offering a full-service AI infrastructure marks a significant strategic evolution.

Meta’s partnership with Cerebras Systems could significantly reshape AI development. For an instance, with over 2,600 tokens per second, this collaboration enables real-time AI applications that were previously impractical. Developers can now build low-latency voice assistants, interactive code generation tools, and instant multi-step reasoning systems.

Traditional AI inference relies heavily on GPUs, but Cerebras’ Wafer-Scale Engine offers an alternative that could challenge Nvidia’s dominance in AI hardware. This shift might encourage more companies to explore custom AI chips for efficiency gains.

For an uninitiated, AI inference is the process where a trained AI model applies its learned knowledge to make predictions or decisions on new data. It’s essentially the "thinking" phase of AI—where it takes what it learned during training and uses it in real-world applications.

By integrating Cerebras’ speed into the Llama API, Meta is making high-performance AI more accessible to developers worldwide. This could accelerate innovation across industries, from quick commerce automation to climate modeling—areas you’ve explored extensively.

Andrew Feldman, CEO and co-founder of Cerebras, said, “Cerebras is proud to make Llama API the fastest inference API in the world. Developers building agentic and real-time apps need speed. With Cerebras on Llama API, they can build AI systems that are fundamentally out of reach for leading GPU-based inference clouds.”

Cerebras is the fastest AI inference solution as measured by third party benchmarking site Artificial Analysis, reaching over 2,600 token/s for Llama 4 Scout compared to ChatGPT at ~130 tokens/sec and DeepSeek at ~25 tokens/sec.

Accenture Launches Accenture AI Refinery™ Framework Built on Nvidia AI Foundry, Also Introduces Llama 3.1 Collection of Openly Available Models

Accenture Launches Accenture AI Refinery™ Framework Built on Nvidia AI Foundry, Also Introduces Llama 3.1 Collection of Openly Available Models

Accenture has just introduced the Accenture AI Refinery™ framework, which leverages NVIDIA AI Foundry. This framework empowers clients to create custom Llama 3.1 language models. These models can be trained on enterprise-specific data and tailored to address unique business needs. By using generative AI, organizations can drive reinvention and transform their industry.

Accenture also launched Llama 3.1 collection refers to a set of openly available language models. These models are part of the Accenture AI Refinery™ framework, which Accenture recently launched. Organizations can use these models as a foundation to build custom language models tailored to their specific needs. By refining and training these prebuilt models with proprietary data, businesses can create powerful AI solutions that address unique challenges in their industry.

The AI Refinery framework includes four key elements:

1. Domain Model Customization and Training: Refining prebuilt foundation models with proprietary data and processes.

2. Switchboard Platform: Allows users to select model combinations based on context, cost, or accuracy.

3. Enterprise Cognitive Brain: Indexes corporate data and knowledge for gen-AI applications.

4. Agentic Architecture: Enables autonomous AI systems to reason, plan, and propose tasks.

These services will be available to all customers using Llama in Accenture AI Refinery, which is built on the NVIDIA AI Foundry service comprised of foundation models, NVIDIA NeMo and other enterprise software, accelerated computing, expert support and a broad partner ecosystem. Models created with AI Refinery can be deployed across all hyper scaler clouds with a variety of commercial options.

This development marks a significant step forward in enterprise generative AI adoption, allowing businesses to create and deploy custom models that align with their unique priorities and industry requirements.

Julie Sweet, chair and CEO of Accenture, said, “The world’s leading enterprises are looking to reinvent with tech, data and AI. They see how generative AI is transforming every industry and are eager to deploy applications powered by custom models. Accenture has been working with NVIDIA technology to reinvent enterprise functions and now can help clients quickly create and deploy their own custom Llama models to power transformative AI applications for their own business priorities.”

Jensen Huang, founder and CEO of NVIDIA, said, “The introduction of Meta’s openly available Llama models marks a pivotal moment for enterprise generative AI adoption, and many are seeking expert guidance and resources to create their own custom Llama LLMs. Powered by NVIDIA AI Foundry, Accenture’s AI Refinery will help fuel business growth with end-to-end generative AI services for developing and deploying custom models.”

Accenture is also using this AI Refinery framework to reinvent its enterprise functions, initially with marketing and communications and then extending to other functions. The solution is enabling Accenture to quickly create gen AI applications that are trained for its unique business needs.

Market Reports

Market Report & Surveys
IndianWeb2.com © all rights reserved