ScaleOps’ cover photo
ScaleOps

ScaleOps

Software Development

New York, NY 22,536 followers

Autonomous cloud & AI resource management. Built for production. Trusted by the world's leading companies.

About us

ScaleOps is redefining cloud resource management from the ground up. Led by a team of cloud infrastructure experts and built for critical production environments, ScaleOps is on a mission to build the Cloud Operating System for the AI era, one that unlocks efficiency and scale while maximizing performance in critical and complex production environments. By bringing real-time, application context-aware automation to cloud resource management, the ScaleOps platform helps organizations eliminate waste, reduce costs, and run critical applications with confidence across any environment.

Website
https://www.scaleops.com
Industry
Software Development
Company size
51-200 employees
Headquarters
New York, NY
Type
Privately Held
Founded
2022
Specialties
Kubernetes, Cloud Infrastructure, Resource Optimization, Cost Reduction, DevOps, FinOps, Continuous Optimization , and Cost Optimization

Products

Locations

Employees at ScaleOps

Updates

  • 90 seconds is not enough to show what a team of A-Players this is. We tried anyway. ScaleOps just wrapped its annual Scalers Club abroad - a weekend away from the day-to-day to recharge, bond, and enjoy being together. We're building something big at ScaleOps. Weekends like this are a reminder of what makes it possible - teams who actually like spending time together build better, bigger things together. And we're just getting started. Want to be on the next trip? We're hiring across all teams: https://lnkd.in/deVU_p64 Yassou 🇬🇷

  • ScaleOps automates Kubernetes resource management. Over-provisioned pods, idle GPU memory, replicas nobody needs. Most teams fix this by hand, sometimes, when someone has time. ScaleOps does it autonomously instead: Real-Time Pod Rightsizing, Smart Pod Placement, and Node Optimization, running continuously in production. Less manual tuning, more consistent performance, and lower cloud costs as a result.

  • Almost every Kubernetes cluster has at least one pod that can never be evicted, and most teams do not know it is there. New video: the 3 patterns that quietly block Karpenter and Cluster Autoscaler, and how to check for them. Watch below ⬇️

  • Always-on AI agents have a real cost. Most of it is a GPU sitting idle behind them, 77% of the time on the fleet we measured. Horizontal Pod Autoscaler scales on CPU and memory by design. It has no native GPU signal, so it once collapsed this fleet to one replica while the GPU sat pegged at 100%. Read the blog to see how we killed the idle GPU burn on this fleet, cold starts and all, without adding latency. Link in comments ⬇️

    • No alternative text description for this image
  • ScaleOps reposted this

    Last month, we wrapped up our strongest quarter to date. Celebrating growth and records is the easy part. The number that tells the real story is ARR per Scaler. We grew the team a lot over the last 12 months, and ARR per Scaler still grew 1.95x. That says something about the market, the product, and mostly - about the people building ScaleOps. AI workloads are everywhere now - inference, LLMs, AI agents, MCP servers, tools. All of it requires heavy compute, and all of it is very bursty. Companies cannot manage their resources manually at that scale. This is creating massive demand for autonomous AI infrastructure management, the problem we have spent four years building for. It drove our growth over the last few quarters, and demand is bigger than ever. We run AI-first as a principle. We research with AI, build with AI, provide value to our customers with AI, and decide with AI. And we have A players who are customer obsessed. AI makes a great team faster. That combination is the whole story. Proud to be building ScaleOps in this era, with this team. Nir Cohen, Taylor Grabus, Zak Blawie ☁☸️, Yarden Weber, Ben Grady, Joey Balázs, Rotem Ben Hamou, Eyal Zilberberg, Adi Steiner, and all the Scalers! This is just the beginning 🚀

    • No alternative text description for this image
  • If your GPU says it's 90% utilized, it's probably lying to you. Why? That number is SM utilization, sampled at an instant. nvidia-smi catches the GPU mid-compute and reports it busy. But if the training or inference loop keeps stalling to load the next batch, the time-averaged utilization can sit far lower than the dashboard suggests. You're basically reading a snapshot and paying for the whole hour. This is why so many production inference workloads run at 5 to 20% real GPU utilization while the bill charges for 100% of the device. Kubernetes hands out GPUs as whole units, but inference consumes compute and memory unevenly, so the gap between allocated and actually used is where the money is wasted. The fix starts with measuring the right thing: SM utilization over time and framebuffer memory, not a single nvidia-smi glance. We published an article explaining how to measure real GPU demand and close that gap. Go check it out! Link in comments 👇

  • Manual GPU tuning doesn't scale with AI workloads that change by the minute. ScaleOps autonomously observes how each AI workload actually uses its GPU, compute and memory, in real time. From there, it automatically assigns the right fractional GPU policy and keeps adjusting allocation as usage shifts. No static slicing. No manual retuning. Just GPUs running at the density your workloads actually need.

Similar pages

Browse jobs