Your company needed agents ...
self-assessment
You built agents. Mission accomplished?
You put Claude behind Slack, skills on cron, and a memory DB you designed. It works, and your team uses it (right?). Now what?
See how you stack up
Marco
RevOps lead, 150-person cloud infrastructure company
38 / 100A scheduler with good prompts
# memory
Eight days: Claude behind Slack, memory in a 23-table SQLite database. His CRO called it the best internal tool the company had shipped. Then the team passed a hundred people, the AE coaching rubrics did not fit CS or TAMs, and the bot began calling itself a "CRO bot" because it had absorbed Marco's personality. Per-user personalization turned out to be a concept his architecture did not have.
The system worked for Marco. The day he tried to make it work for everyone else, he found he had built a different product than the one his team needed.
Priya
GTM engineering lead, fast-growing AI-native startup
44 / 100A system, not yet a loop
# measurement
Skills on cron, event triggers off the call recorder, state in a homegrown database. It worked. Then her VP of Sales asked how much time it had given the team back, and whether calendars looked different from six months ago. She had nothing: no engagement data, no before-and-after, no idea which skills people read versus ignored. The bot's weekly self-report was the bot grading its own homework.
Priya had built a system she could not measure, so she could not tell leadership whether it was working, or tell herself whether it was drifting.
Dex
Technical co-founder, developer tools company
31 / 100A scheduler with good prompts
# talk-back
His team wired up exactly what he thought an agentic CRM was: Claude compute on a Slack event or a cron job. Then a senior AE replied to the bot, "stop including the demo pipeline in my briefing." The message went nowhere. By the third week the AE had stopped reading, and so had three other reps. Dex heard about it secondhand, from a manager.
Dex's users were talking to a wall. Every other AI tool in their lives listens, and when this one did not, nobody filed a ticket. They just stopped showing up.
Anika
Head of operations, Series B infrastructure company
47 / 100A system, not yet a loop
# harvest
Every call graded against a methodology rubric in Confluence, feedback in Slack within fifteen minutes. Six months in, her best rep, Tomás, was quietly ignoring it: "It keeps telling me to do discovery differently, but my way is closing deals." He was right. He had evolved past the rubric, the rubric had no way to learn from him, and her newest reps got the same coaching he did.
Anika's coaching was frozen in time. Her best rep's breakthroughs stayed trapped with him, and the rubric was already behind the market.
Ravi
Head of revenue operations, cybersecurity company
29 / 100A scheduler with good prompts
# permissions
Deal intelligence, competitive alerts, pipeline analysis, all through one shared Slack bot. His CISO had one question: who can see what? The bot ran on a single service account and posted summaries into a shared channel, surfacing details not everyone there was cleared to see. Fixing it had been on Ravi's list for two months.
Ravi's system had no permission model. Nothing in the architecture kept each person's agents to what that person was allowed to see.
Suki
GTM engineer, product-led growth company
36 / 100A scheduler with good prompts
# authorship
Twelve skills, each prompt hand-tuned over weeks. Then the company hired fifteen reps in Q3 and her manager asked for personalized coaching for each. A day per skill meant the first would be stale before the fifteenth shipped. She also noticed her twelve existing prompts were byte-identical to the day she wrote them, a thousand runs later. The model underneath had improved; her system had not.
Improvements needed her keyboard, new people her calendar. She was not running a system. She was the system.
Nate
Sales operations lead, late-stage startup
42 / 100A system, not yet a loop
# governance
Deal analysis, pipeline reviews, call coaching, competitive intel: one bot, one personality, one set of prompts. Then his VP of Sales asked to lock down the forecast methodology skill while letting reps customize their coaching. Nate's system had one mode, everything controlled by him in a repo only he could edit. Open access lost the forecast; locking it lost personalization.
Nate had governance by accident. His system could not tell "nobody touches this" from "make it your own," so it did neither well.
Leah
RevOps leader, mid-market SaaS company
34 / 100A scheduler with good prompts
# loops
Skills built, tools connected, team deployed. Three months in she took a two-week vacation and came back to a system that had not moved. Prompts identical, competitive intel citing pricing from two quarters ago, a product launch none of the skills knew about. Her team had stopped reading the output weeks earlier, and nobody had told her.
Leah's system did not compound. Its improvement rate was the free hours of one person, and that person had been on a beach.
# 01
The remaining gaps
# 02
How the grading works
Ten multiple-choice questions. Each offers four self-descriptions; pick the one that is uncomfortably, recognizably you. You get a score out of 100, a tier, and a per-question breakdown with what good looks like beside each answer. About four minutes. No signup for the score. The full report asks for a work email.