Lambda’s cover photo
Lambda

Lambda

Software Development

San Francisco, California 58,255 followers

The Superintelligence Cloud

About us

The Superintelligence Cloud

Website
http://lambda.ai/linkedin
Industry
Software Development
Company size
501-1,000 employees
Headquarters
San Francisco, California
Type
Privately Held
Founded
2012
Specialties
Deep Learning, Machine Learning, Artificial Intelligence, LLMs, Generative AI, Foundation Models, GPUs, Distributed Training, Superintelligence, AI Infrastructure, and AI Factories

Locations

Employees at Lambda

Updates

  • View organization page for Lambda

    58,255 followers

    GPU utilization increased from ~20% to 43% on a reservation of 96 NVIDIA H100 GPUs, while cutting queue starvation by 74%, with blocked jobs falling from roughly 10 per day to around 4. That’s the concrete result SPREEAI saw after fixing their orchestration. When a unified diffusion model requires 80–100 GB of memory, you can’t simply throw workloads at a cluster and expect to use those GPUs efficiently. SPREEAI was dealing with workload fragmentation, ad-hoc submissions, and storage I/O blocking that left expensive GPUs idle. Working with Lambda’s ML engineering team, they implemented MLflow-based experiment orchestration with structured queuing and workload matching. They also connected Lambda’s Prometheus APIs to Grafana for real-time visibility into utilization gaps. The video testimonial covers how they diagnosed the bottlenecks and what the remediation looked like.

  • View organization page for Lambda

    58,255 followers

    Evaluating image editing models with a single score hides the nuance behind a lower score. EdiVal-Agent (ICLR 2026) splits evaluation across instruction following, content consistency, and visual quality. Its agentic judge reached 81.3% agreement with human judgments, compared with 75.2% for a VLM-only evaluator and 68.9% for CLIP_dir. The failure that stands out is how quickly multi-turn editing exposes weak preservation. Each new instruction has to land without undoing earlier edits or changing unrelated content. Errors compound even when single-turn results look strong. EdiVal-IF measures instruction following. EdiVal-CC covers content consistency, while EdiVal-VQ checks visual quality. https://lnkd.in/eVidPybM

    • Bar chart showing agreement percentages across various image editing tasks, comparing different evaluation methods and their scores.
    • No alternative text description for this image
  • View organization page for Lambda

    58,255 followers

    Build or buy AI? Wrong question. Lambda's Robert Brooks IV is on the main stage at AI Infra Summit 2026 with JLL CTO Yao Morin, U.S. Bank EVP & Chief AI Officer Prashant Mehrotra, and Carrier Chief Data & AI Officer Arun Nandi, and the room's landing on the same answer: it's build AND buy. New term coined: valuemaxxing. Using value per task as a lens to govern resource allocation, while architecting around agility (easy model swaps) and availability (SLAs and model lifecycle).

    • No alternative text description for this image
  • View organization page for Lambda

    58,255 followers

    MLPerf Inference v6.1 is out, and we have two firsts: the first agentic inference workload on datacenter hardware, and the first MLPerf deployment of a model over a trillion parameters. On 4x NVIDIA Blackwell Ultra GPUs, we posted the leading Offline throughput on GPT-OSS 120B among all Blackwell Ultra GPU submissions, plus an 8.85% throughput gain over v6.0 on identical hardware. That's six months of pure software optimization. On NVIDIA HGX B200, we swapped in Kimi K2.6 for the open-division agentic benchmark, running a 1T+ parameter model where the reference workload expects 27B. Same harness, no memory ceiling. VLM inference made its debut in our lineup too. Qwen3-VL on NVIDIA HGX B200 came in 28.5% faster (Offline) than the fastest v6.0 submission on comparable hardware. The agentic number that matters most: 1,007 of 1,007 replay turns completed, zero failures, 86.83% BFCL v4 accuracy, at a fraction of the latency of the 27B reference deployments. Full results and methodology in the blog. Link below. https://lnkd.in/e8udxxeM

    • No alternative text description for this image