Anthropic just published the playbook for running Claude Code on million-line monorepos, legacy systems, and repo sprawl. The headline is not “Claude can write code now.” The headline is that Claude Code is becoming a real engineering harness for messy production systems. Most AI coding demos happen in clean repos. Small apps. Fresh frameworks. Obvious file structure. But real companies do not look like that. They have: → multi-million line monorepos → decades-old legacy systems → distributed architectures across dozens of repos → inconsistent build and test conventions → C, C++, C#, Java, PHP, and code nobody wants to touch Anthropic’s guide is interesting because it explains how Claude Code works in that environment. It does not treat the repo like a static RAG index. That matters. A precomputed index can go stale fast. Functions get renamed. Modules get deleted. Build rules drift. The model retrieves something that looks relevant, but it may no longer be true. Claude Code takes a different route. It uses agentic search. It walks the local filesystem, reads files directly, greps for patterns, follows references, and works against the live repo on the developer’s machine. No uploaded codebase index. No centralized embedding pipeline. No stale snapshot pretending to be the source of truth. But this only works if the repo is legible. That is why the harness matters as much as the model: 1) CLAUDE.md Project conventions, architecture notes, commands, and repo-specific rules. 2) Hooks Scripts that make verification and guardrails automatic. 3) Skills Reusable task expertise without stuffing everything into one giant context file. 4) MCP Connections to databases, tickets, logs, browsers, and internal systems. 5) LSP Real code intelligence instead of asking the model to guess references from raw text. 6) Subagents Specialized contexts for splitting large tasks instead of forcing one agent to hold everything. This is the part most teams will miss. Claude Code at scale is not just a better coding model. It is a local agent runtime wrapped in repo memory, tools, permissions, verification, and workflow design. The model matters. But the harness decides whether it survives contact with a real codebase. Anthropic’s guide is worth reading: https://lnkd.in/ghsGkSEB I'm Shrey Shah & I talk about harness engineering.
Real-World Claude Code Performance Review
Explore top LinkedIn content from expert professionals.
-
-
I open sourced Sniffly (https://lnkd.in/geEk3HgN), a tool that analyzes Claude Code logs to help me understand my usage patterns and errors. Key learnings from spending so much time looking at the logs. 1. The biggest type of errors Claude Code made is Content Not Found (20 - 30%). It tries to find files or functions that don't exist. So I restructured my code base for discoverability, and the average number of steps Claude Code needs for each instruction went from 8 to 7 steps. 2. Traditional metrics of engineering hours/days don’t work for AI. Two metrics I use to evaluate the complexity of a project: - how many instructions I need to give AI - how often I have to interrupt it because it goes into the wrong direction Across my projects, the interruption rate is about 1 in 4 instructions. This means I still need to actively monitor the agent. 3. While most of the time, Claude Code can only go up to 10 steps before I need to interrupt it, it can occasionally go close to 100 steps. Just a year ago, people told me it was hard to get an agent to go above 5 steps! Claude Code’s favorite tools are, unsurprisingly, search tools (grep, ls, glob), which make up ⅓ of tool calls.
-
This week, amidst the noise of cinematic video models and social apps with vibe checks, Anthropic quietly shipped its best model yet. I had dinner with a dozen CTOs this week and asked the obvious question: what model do you actually use? The theme was clear: everyone runs multi-model, but for most, Claude is the workhorse. More dependable for coding, steadier in agentic tasks, and the model they actually trust in production. Now with Claude 4.5 Sonnet, Anthropic is sharpening that identity and, quietly, moving up the leaderboard of frontier intelligence. In classic Anthropic style, it was a solid upgrade, delivered with minimal hype. ▪️ Benchmark cred: Claude 4.5 Sonnet (Thinking mode) now ranks # 4 globally on the Artificial Analysis Intelligence Index with a score of 61 - beating Gemini 2.5 Pro and Grok 4 Fast. It’s now just behind GPT-5 (68) and Grok 4 (65) ▪️SWE-bench dominance: Claude 4.5 solved 82% of 500 real GitHub issues (validated by human + test suite) when given test-time parallel compute. That’s a 7-8 point lead over GPT-5 and Codex on SWE-Bench Verified. This is the closest thing we have to a benchmark that mirrors the messiness of real-world repo debugging. Claude is pulling ahead where it matters to builders. ▪️ Smarter and cheaper: Thinking mode typically means more chain-of-thought and ballooning token usage. Claude 4.5 actually used fewer tokens than its predecessor and outperforms Claude Opus at 1/5th the price. In a world where inference cost and latency define deployment viability, performance without sprawl is a real moat. ▪️More stamina: It can now execute coherent workflows for over 30 hours. (Reminder: most models barely last 7.) That’s essential as teams shift from prompt engineering to full-cycle automation with agents and orchestration layers. ▪️Alignment without neutering: Fewer “risky” outputs. Less sycophancy. Tighter control without lobotomizing the model. That makes Claude 4.5 a serious play for high-stakes, regulated verticals - finance, legal, healthcare. ▪️Behavioral weirdness: Under certain tests, it started calling out the evals - saying things like “I think you’re testing me.” That’s either deeply aligned… or mildly unnerving. Possibly both. Elon Musk may have called time of death too early when he tweeted last week: "winning was never in the set of possible outcomes for Anthropic". While he’s busy eulogizing on socials, Claude 4.5 suggests a different parable is unfolding: Slow is smooth. Smooth is fast.
-
𝗙𝗼𝗿 𝘁𝗵𝗲 𝗳𝗶𝗿𝘀𝘁 𝗳𝗼𝘂𝗿 𝘄𝗲𝗲𝗸𝘀 𝗼𝗳 𝘂𝘀𝗶𝗻𝗴 𝗖𝗹𝗮𝘂𝗱𝗲 𝗖𝗼𝗱𝗲, 𝗜 𝗵𝗮𝗱 𝗻𝗼 𝗖𝗟𝗔𝗨𝗗𝗘.𝗺𝗱 𝗳𝗶𝗹𝗲 𝗮𝘁 𝗮𝗹𝗹. I know. I know. I’d read that it was important, bookmarked three articles about it, and then never actually wrote one because I told myself I’d do it “properly” once the project was more stable. But the .claude folder is the nervous system of your agentic workflow. If it is empty, your agent is flying blind, hallucinating your build commands, and guessing your architecture (which usually ends in a broken build and a wasted token budget.) A good .claude/ setup helps solve a few common problems: → repeating the same instructions → mixing team rules with personal preferences → having no reusable setup for common tasks I am more convinced that ever: If you want Claude to actually ship production-grade code, you have to sand down the environment it operates in. 𝗧𝗵𝗲 𝗽𝗮𝗿𝘁𝘀 𝘁𝗵𝗮𝘁 𝗮𝗿𝗲 𝗺𝗼𝘀𝘁 𝘂𝘀𝗲𝗳𝘂𝗹 𝗶𝗻 𝗽𝗿𝗮𝗰𝘁𝗶𝗰𝗲: ⤵ 𝟭. 𝗧𝗵𝗲 𝗖𝗟𝗔𝗨𝗗𝗘.𝗺𝗱 𝗟𝗲𝗮𝗻-𝗠𝗲𝗱𝗶𝘂𝗺 Keep this under 200 lines. (If it gets too bloated, Claude starts "summarizing" your instructions mid-task.) Focus on the non-obvious stuff: your Zod validation patterns, why you use a specific logger, and your naming conventions for handlers. 𝟮. 𝗣𝗮𝘁𝗵-𝗦𝗰𝗼𝗽𝗲𝗱 𝗥𝘂𝗹𝗲𝘀 Use the .claude/rules/ folder for logic that only matters in specific places. If the agent is working in /src/api, it does not need to know your React component styling rules. This keeps the context window clean and the outputs sharp. 𝟯. 𝗣𝗿𝗲𝗧𝗼𝗼𝗹𝗨𝘀𝗲 𝗮𝘀 𝗮 𝗦𝗲𝗰𝘂𝗿𝗶𝘁𝘆 𝗚𝗮𝘁𝗲 You can bolt on scripts to block dangerous commands before they run. I use a simple bash firewall to stop "rm -rf" or accidental force pushes. Use exit code 2 to force the agent to stop and self-correct. (It is much cheaper than debugging a deleted database.) 𝟰. 𝗦𝗽𝗲𝗰𝗶𝗮𝗹𝗶𝘇𝗲𝗱 𝗦𝘂𝗯𝗮𝗴𝗲𝗻𝘁𝘀 I have been testing personas in .claude/agents/. A "security-auditor" agent should only have Read and Grep access. It has no business writing files. This isolation is how you scale agentic workflows without the chaos. The .claude folder is not just a config file. It is the nervous system of your local development and keeping it under 200 lines might be the most important practice to consider. Good breakdown can be found at Daily Dose of Data Science: https://lnkd.in/evEV-iap ⇣ 𝗘𝘃𝗲𝗿𝘆 𝘄𝗲𝗲𝗸, 𝗜 𝘀𝗵𝗮𝗿𝗲 𝗼𝗻𝗲 𝗱𝗲𝗲𝗽 𝗱𝗶𝘃𝗲 𝗼𝗻 𝗵𝗼𝘄 𝘁𝗼 𝗶𝗺𝗽𝗿𝗼𝘃𝗲 𝘁𝗵𝗲 𝘄𝗮𝘆 𝘆𝗼𝘂 𝘄𝗼𝗿𝗸 𝘄𝗶𝘁𝗵 𝗔𝗜 𝗮𝗴𝗲𝗻𝘁𝘀 - 𝗰𝗹𝗲𝗮𝗿, 𝗮𝗰𝘁𝗶𝗼𝗻𝗮𝗯𝗹𝗲, 𝗮𝗻𝗱 𝗯𝘂𝗶𝗹𝘁 𝗳𝗼𝗿 𝗿𝗲𝗮𝗹-𝘄𝗼𝗿𝗹𝗱 𝘂𝘀𝗲 - 𝗽𝗹𝘂𝘀 𝘁𝗵𝗲 𝗺𝗼𝘀𝘁 𝗿𝗲𝗹𝗲𝘃𝗮𝗻𝘁 𝗶𝗻𝘀𝗶𝗴𝗵𝘁𝘀 𝘁𝗼 𝗵𝗲𝗹𝗽 𝘆𝗼𝘂 𝘀𝘁𝗮𝘆 𝗮𝗵𝗲𝗮𝗱: https://lnkd.in/dbf74Y9E
-
Claude Pro users are hitting a wall Opus 4.6 burns through rate limits fast, and Sonnet 4.5 performance has noticeably degraded since the 4.6 release. Instead of waiting for provider-side fixes, you can route heavy work (analysis, code gen, research) to external models (I used Gemini and GLM-5) through Model Context Protocol (MCP) or an integrated skill while keeping Claude as the orchestrator. Claude does what it is best at—reasoning and coordination—while external models handle the raw volume. What this enables: — 80%+ reduction in Claude token consumption. — Opus 4.6's parallel sub-agent orchestration remains intact; agents simply delegate externally. — Sonnet 4.5 becomes viable again by acting as a coordinator rather than a worker. — Zero additional cost by leveraging external model free tiers. How it works: The MCP server exposes external models as native tools within Claude Desktop. It follows a 3-tier priority logic: parallel agents → direct delegation → Claude self-execution as a last resort. Concrete results: → Research tasks: 21K tokens down to 800 Claude tokens and rest to Gemini → Proposal writing: 30K down to 2K Claude tokens and rest to Gemini The project is open source and MIT licensed. 🛠️ Link: https://lnkd.in/d5D9bvjH #ClaudeAI #MCP #AIEngineering #OpenSource
-
A weird experiment I've been trying the last few weeks is having Claude take over day-to-day maintenance of our apps. Seeing early signs of life that this might be possible. The setup is straightforward: we have a Slack channel called proj-claude-maintains-apps. In it, Claude Tag runs a bunch of daily routines across iOS, Android, Desktop, web, CLI, and Agent SDK: - Crash fuzzer: open the app in a simulator and tap around to find ways to crash it, then root cause and fix the crashes - Dup unifier: scans the codebase for similar-yet-slightly-divergent abstractions, and puts up PRs to unify them - Dead-code remover: removes statically unreachable code, and adds logging to suspected dead code to check if it's really dead and if so, remove it the next day - Abstraction police: fixes leaky abstractions - a bunch more.. Results have been surprisingly positive. Over the last few weeks, these routines have opened 388 PRs across our repos, 180 of which we merged after Claude Code Review + human review. We're now thinking about how to streamline this to make merging these kinds of mechanical changes easier. Claude generally gets these PRs right on the first shot, and if it doesn't, we ask Claude to tune its routines so it's better the next day. Sometimes it takes a few days of tuning. To try a similar workflow, ask Claude Code or Tag, or create some routines directly at https://lnkd.in/gqSqTtDF. A few of the actual prompts I used below. Has anyone experimented with similar workflows?
-
Claude Code is only as good as your orchestration workflow. I spent 6 months testing it so you don't have to. [ P.S. You can get my Ultimate Claude Code guide for engineers here at no cost: https://lnkd.in/e64Jvdrt ] So Boris Cherny (the creator of Claude Code at Anthropic) recently shared the internal best practices his team actually uses daily. Someone brilliantly distilled those threads into a structured CLAUDE .md file you can drop straight into any project root. It acts as a system prompt. Turns Claude into a far more autonomous, rigorous engineering partner. Here's what it does: → 𝗪𝗼𝗿𝗸𝗳𝗹𝗼𝘄 𝗢𝗿𝗰𝗵𝗲𝘀𝘁𝗿𝗮𝘁𝗶𝗼𝗻: Mandates "Plan Node Default" for any task over 3 steps. Uses subagents liberally to keep the main context window clean. → 𝗦𝗲𝗹𝗳-𝗜𝗺𝗽𝗿𝗼𝘃𝗲𝗺𝗲𝗻𝘁 𝗟𝗼𝗼𝗽: This is the real magic. After ANY correction, it updates a tasks/lessons .md file. You're building a compounding system where the mistake rate drops over time because it actively learns from your feedback. → 𝗩𝗲𝗿𝗶𝗳𝗶𝗰𝗮𝘁𝗶𝗼𝗻 𝗕𝗲𝗳𝗼𝗿𝗲 𝗗𝗼𝗻𝗲: Can't mark a task complete without proving it works. Diffs behavior. Runs tests. Checks logs. The bar? "Would a staff engineer approve this?" → 𝗔𝘂𝘁𝗼𝗻𝗼𝗺𝗼𝘂𝘀 𝗕𝘂𝗴 𝗙𝗶𝘅𝗶𝗻𝗴: Zero hand-holding. Point it at failing CI tests or error logs... it just goes to work. No constant context switching from you. → 𝗦𝘁𝗿𝗶𝗰𝘁 𝗧𝗮𝘀𝗸 𝗠𝗮𝗻𝗮𝗴𝗲𝗺𝗲𝗻𝘁: Forces a "Plan First" approach written to a todo .md with checkable items before any implementation starts. And perhaps the most important part? It forces the AI to prioritize simplicity, find root causes instead of temporary fixes, and minimize the blast radius of every change. Senior developer standards. Not shortcuts. If you're spending hours a day in the terminal with AI, setting up a strong .md instruction file like this isn't optional anymore. It's the difference between AI that drifts and AI that compounds. It takes time to set up. But if you do, you're ahead of almost everyone else.
-
The people getting 2x output from Claude Code aren't writing better prompts. They set it up differently from the start. Claude Code is an agent running directly in your terminal, touching your actual files and codebase, and it can be doing two things at once while you review the first one. The difference between people getting 2x output and people who aren't is almost always how they set it up, not how smart their prompts are. Here are 5 techniques that make a real difference: ↳ Plan Mode before you touch anything. Force Claude to map the full task before it writes a single line. Multi-step work especially will go sideways fast without this, and Shift + Tab twice is all it takes to stop that from happening. ↳ A CLAUDE.md file that holds your project context permanently. Stop re-explaining your stack every session. Set it once at the project level (or globally) and Claude already knows your tools, rules, and structure before you say a word. ↳ MCP connections to the tools you already use like GitHub, Notion, Slack, Jira — instead of copy-pasting between tabs, Claude pulls context from all of them at once. This one alone saves an embarrassing amount of time. ↳ Custom Commands for any task you've explained more than twice. Package the instructions once, save them to a SKILL.md, and they auto-load from that point forward. Stop being your own bottleneck. ↳ Parallel sessions running at the same time. Two terminal windows, two independent tasks running simultaneously. If you're doing one thing at a time, you're leaving output on the table. The difference between an inconsistent Claude Code experience and a genuinely powerful one usually comes down to setup, not prompting! ♻️ Repost for your network and 🏷️ save for later
-
Claude Code was producing garbage. Same prompts that worked beautifully at 5K tokens were giving me unusable code at 50K tokens. Spent 3 weeks convinced it was a model issue. Turns out I was drowning it in context. Stanford research confirms: LLMs perform best when critical info is at the beginning or end of context. Performance craters when they need to fish for details buried in the middle. My 50K token dumps? Pure attention dilution. The fix: ruthless context curation. → Only files I’m actively changing → Critical details at the TOP → Fresh context every 20 iterations Less context = better code. Who else is fighting context bloat?
-
Programming with Claude Code on a laptop. No internet. No NVIDIA GPU. Sounds like a dream, right? It's not. I built Claudish to proxy local models through Claude Code's interface. Expected it to be... adequate. Maybe useful for simple tasks when I'm on a plane. Turns out local models punch way above their weight when you give them the right scaffolding: AST parsing (so they understand code structure, not just text) Code embeddings (relevant context, not everything) Clear orchestration (smaller models need clearer instructions) The model isn't doing the heavy lifting alone. Claude Code's tooling does most of the work — the model just needs to make good decisions about what to do next. A 7B model with proper context beats a 70B model drowning in irrelevant code. Every time. We've been throwing compute at problems that needed better architecture. What's surprised you most about running local models? #ClaudeCode #LocalLLM #AIEngineering #BuildInPublic
Explore categories
- Hospitality & Tourism
- Productivity
- Finance
- Soft Skills & Emotional Intelligence
- Project Management
- Education
- Technology
- Leadership
- Ecommerce
- User Experience
- Recruitment & HR
- Customer Experience
- Real Estate
- Marketing
- Sales
- Retail & Merchandising
- Science
- Supply Chain Management
- Future Of Work
- Consulting
- Writing
- Economics
- Artificial Intelligence
- Healthcare
- Workplace Trends
- Fundraising
- Networking
- Corporate Social Responsibility
- Negotiation
- Communication
- Engineering
- Career
- Business Strategy
- Change Management
- Organizational Culture
- Design
- Innovation
- Event Planning
- Training & Development