SMALL LANGUAGE MODELS: TAKING AI TO THE EDGE

SUMIT SHARMA

Session Overview

From Model Overload to Small, Local, and Multi-Agent AI.

Sumit Sharma opened by sizing up just how fast Generative AI is moving — with over 378,000 language models listed on Hugging Face, and a single 10-day window in which Microsoft, Meta, Databricks, and Google all shipped major new models (Phi-3, LLaMA 3, DBRX, and Gemini 1.5 Pro respectively). That pace, he argued, is exactly why enterprises struggle: in 2023 only about 10% of Generative AI pilots made it to production, largely because organizations lacked "AI-ready data," leading to data drift, high infrastructure costs, and frustrating latency once real enterprise data entered the picture.

He walked through the "build vs. buy" decision enterprises face — SaaS models like ChatGPT are easy to adopt but behave as a black box, which doesn't sit well with AI ethics officers who need to know what data a model was trained on. This sets up the core theme of the talk: the shift toward Small Language Models (SLMs) that run at the edge — directly on factory floor machines, IoT devices, or laptops — which matters enormously for regulated sectors like defense and security that can't send data to the cloud.

Using Microsoft's Phi-3 as a concrete example, he showed how a 3.8-billion-parameter model, focused on data quality over data quantity and shrunk via quantization to fit in just 2GB of RAM, can run locally on a device like an iPhone 14 — while still offering a 4K context window extendable to 128K via "long rope." He closed by describing the road ahead: rather than one giant LLM doing everything, enterprises will orchestrate multiple specialized SLMs (augmented with RAG for missing knowledge) in multi-agent systems — and leaders need to treat Generative AI as a long-term strategic investment, not judge it on one or two short-term use cases, with the end goal being augmented humans freed up for high-value work.


Key Takeaways & Concepts

  • An Overwhelming Pace: 378,000+ models on Hugging Face, with Phi-3, LLaMA 3, DBRX, and Gemini 1.5 Pro all launching within a 10-day window — making model selection a real enterprise challenge.
  • Why Pilots Fail: Only ~10% of 2023 Generative AI pilots reached production, largely due to a lack of "AI-ready data" causing data drift, plus high infrastructure costs and latency.
  • Build vs. Buy: SaaS models like ChatGPT are easy but operate as a black box — a problem for AI ethics officers who need training-data transparency.
  • The Shift to SLMs & Edge AI: Small Language Models run locally on factory machines, IoT devices, and laptops — critical for regulated sectors like defense that can't send data to the cloud.
  • Phi-3 in Practice: A 3.8B-parameter model, quantized down to 2GB of RAM, can run locally on an iPhone 14 with a 4K context window extendable to 128K via "long rope."
  • Quality Over Quantity: Modern SLMs are trained on curated, high-quality and synthetic data rather than the entire internet — and gaps in factual knowledge are filled using RAG.
  • Multi-Agent Future: Enterprises will move from one giant LLM to orchestrated multi-agent systems of specialized SLMs working together on complex workflows.
  • Strategic ROI, Human Augmentation: Generative AI should be judged as a long-term strategic investment — the goal is augmenting humans to focus on high-value work, not replacing them.

Session Highlights

Sumit Sharma presenting at AI Dev Day India 2024
Sumit Sharma session moment
Audience engaging with Sumit Sharma's session
Sumit Sharma Q&A

Up Next:

Beyond the Hype

L Venkata Subramaniam's Session