90 seconds is not enough to show what a team of A-Players this is. We tried anyway. ScaleOps just wrapped its annual Scalers Club abroad - a weekend away from the day-to-day to recharge, bond, and enjoy being together. We're building something big at ScaleOps. Weekends like this are a reminder of what makes it possible - teams who actually like spending time together build better, bigger things together. And we're just getting started. Want to be on the next trip? We're hiring across all teams: https://lnkd.in/deVU_p64 Yassou 🇬🇷
ScaleOps
Software Development
New York, NY 22,536 followers
Autonomous cloud & AI resource management. Built for production. Trusted by the world's leading companies.
About us
ScaleOps is redefining cloud resource management from the ground up. Led by a team of cloud infrastructure experts and built for critical production environments, ScaleOps is on a mission to build the Cloud Operating System for the AI era, one that unlocks efficiency and scale while maximizing performance in critical and complex production environments. By bringing real-time, application context-aware automation to cloud resource management, the ScaleOps platform helps organizations eliminate waste, reduce costs, and run critical applications with confidence across any environment.
- Website
-
https://www.scaleops.com
External link for ScaleOps
- Industry
- Software Development
- Company size
- 51-200 employees
- Headquarters
- New York, NY
- Type
- Privately Held
- Founded
- 2022
- Specialties
- Kubernetes, Cloud Infrastructure, Resource Optimization, Cost Reduction, DevOps, FinOps, Continuous Optimization , and Cost Optimization
Products
Automated Kubernetes Optimization
Cloud Management Platforms (CMP)
ScaleOps: Automated Cloud Resource Management for Kubernetes ScaleOps delivers a fully automated platform for managing Kubernetes resources in production. It enables organizations to achieve up to 80% cloud cost savings while maximizing performance or reliability. ScaleOps ensures optimal resource utilization across all your K8s infrastructure.
Locations
-
Primary
Get directions
New York, NY, US
Employees at ScaleOps
Updates
-
We are at the #AWSSummit Tel Aviv 🇮🇱. The ScaleOps team is at Booth G6 ready to show you how we help organizations autonomously manage their AI infrastructure. See you at Booth G6 for a chat!
-
-
ScaleOps reposted this
Big day for ScaleOps 🚀 Gilad Elyashar is joining us as Chief Product Officer. Together with Gilad and all the Scalers, we will continue to build the Autonomous Infrastructure Orchestration Platform for Enterprise AI in Production. Gilad, welcome aboard. Can't wait to build together! 💪
-
-
ScaleOps automates Kubernetes resource management. Over-provisioned pods, idle GPU memory, replicas nobody needs. Most teams fix this by hand, sometimes, when someone has time. ScaleOps does it autonomously instead: Real-Time Pod Rightsizing, Smart Pod Placement, and Node Optimization, running continuously in production. Less manual tuning, more consistent performance, and lower cloud costs as a result.
-
We're live on the ground at ContainerDays in Hamburg today and tomorrow! Come find our team in ScaleOps shirts for a chat. #ContainerDays #Kubernetes
-
-
Always-on AI agents have a real cost. Most of it is a GPU sitting idle behind them, 77% of the time on the fleet we measured. Horizontal Pod Autoscaler scales on CPU and memory by design. It has no native GPU signal, so it once collapsed this fleet to one replica while the GPU sat pegged at 100%. Read the blog to see how we killed the idle GPU burn on this fleet, cold starts and all, without adding latency. Link in comments ⬇️
-
-
ScaleOps reposted this
Last month, we wrapped up our strongest quarter to date. Celebrating growth and records is the easy part. The number that tells the real story is ARR per Scaler. We grew the team a lot over the last 12 months, and ARR per Scaler still grew 1.95x. That says something about the market, the product, and mostly - about the people building ScaleOps. AI workloads are everywhere now - inference, LLMs, AI agents, MCP servers, tools. All of it requires heavy compute, and all of it is very bursty. Companies cannot manage their resources manually at that scale. This is creating massive demand for autonomous AI infrastructure management, the problem we have spent four years building for. It drove our growth over the last few quarters, and demand is bigger than ever. We run AI-first as a principle. We research with AI, build with AI, provide value to our customers with AI, and decide with AI. And we have A players who are customer obsessed. AI makes a great team faster. That combination is the whole story. Proud to be building ScaleOps in this era, with this team. Nir Cohen, Taylor Grabus, Zak Blawie ☁☸️, Yarden Weber, Ben Grady, Joey Balázs, Rotem Ben Hamou, Eyal Zilberberg, Adi Steiner, and all the Scalers! This is just the beginning 🚀
-
-
If your GPU says it's 90% utilized, it's probably lying to you. Why? That number is SM utilization, sampled at an instant. nvidia-smi catches the GPU mid-compute and reports it busy. But if the training or inference loop keeps stalling to load the next batch, the time-averaged utilization can sit far lower than the dashboard suggests. You're basically reading a snapshot and paying for the whole hour. This is why so many production inference workloads run at 5 to 20% real GPU utilization while the bill charges for 100% of the device. Kubernetes hands out GPUs as whole units, but inference consumes compute and memory unevenly, so the gap between allocated and actually used is where the money is wasted. The fix starts with measuring the right thing: SM utilization over time and framebuffer memory, not a single nvidia-smi glance. We published an article explaining how to measure real GPU demand and close that gap. Go check it out! Link in comments 👇
-
Manual GPU tuning doesn't scale with AI workloads that change by the minute. ScaleOps autonomously observes how each AI workload actually uses its GPU, compute and memory, in real time. From there, it automatically assigns the right fractional GPU policy and keeps adjusting allocation as usage shifts. No static slicing. No manual retuning. Just GPUs running at the density your workloads actually need.