返回 Skills
wondelai/skills· MIT 内容可用

high-output-management

Manage for output using Grove''s "High Output Management": a manager''s output is their organization''s output, raised by high-leverage activities. Use when the user mentions "high output management", "managerial leverage", "one-on-ones", "1:1 agenda", "OKRs", "performance review", "task-relevant maturity", "delegation", "meeting overload", "new manager", "how do I run a 1:1", or "just got promoted to manager". Also trigger when structuring a manager''s calendar and meeting cadence, designing team metrics, running planning, coaching delegation, or preparing performance reviews. Covers leverage, production principles, meetings as the medium of management, decisions, OKRs, and task-relevant maturity. For intrinsic motivation, see drive-motivation. For a company operating system, see traction-eos.

安装

与 skills.sh 相同的 Command / Prompt 安装方式


name: high-output-management description: 'Manage for output using Grove''s "High Output Management": a manager''s output is their organization''s output, raised by high-leverage activities. Use when the user mentions "high output management", "managerial leverage", "one-on-ones", "1:1 agenda", "OKRs", "performance review", "task-relevant maturity", "delegation", "meeting overload", "new manager", "how do I run a 1:1", or "just got promoted to manager". Also trigger when structuring a manager''s calendar and meeting cadence, designing team metrics, running planning, coaching delegation, or preparing performance reviews. Covers leverage, production principles, meetings as the medium of management, decisions, OKRs, and task-relevant maturity. For intrinsic motivation, see drive-motivation. For a company operating system, see traction-eos.' license: MIT metadata: author: wondelai version: "1.2.0"

High Output Management

Manage teams the way Andy Grove ran Intel: a manager's output is not what the manager does — it is what their organization produces. This skill turns High Output Management into auditable practice: production principles for knowledge work, output indicators, managerial leverage, meetings as the medium of management, clean decisions, OKRs, and a management style matched to task-relevant maturity.

Core Principle

A manager's output = the output of their organization + the output of the neighboring organizations under their influence. Nothing a manager does — emails, meetings, reviews, decisions — counts in itself; it counts only through how it raises that combined output. Since managerial time is the scarce input, the craft reduces to one question asked relentlessly: of everything I could do right now, what creates the most output per hour spent? Choose high-leverage activities; eliminate negative-leverage ones.

Scoring

Goal: 10/10. Rate management practices, calendars, and processes 0-10 against the principles below. State the current score and the specific changes needed to reach 10/10.

  • 9-10: Output indicators with quality pairs, subordinate-owned 1:1s on a TRM-based cadence, delegation with task-level monitoring, OKRs that stretch without driving pay, calendar built around forecasted key events
  • 7-8: Process meetings run well, but a few activity metrics, ad hoc decision meetings, or skipped training sessions remain
  • 5-6: 1:1s happen irregularly, indicators track busyness, delegation is all-or-nothing, planning produces documents instead of actions
  • 3-4: Management by interruption: status theater, decisions made by rank, reviews as annual surprises
  • 0-2: No 1:1s, no indicators, firefighting as the operating mode, output invisible and unmeasured

Framework

1. Production Principles for Knowledge Work

Core concept: Grove's breakfast factory — deliver a three-minute egg, buttered toast, and hot coffee simultaneously, at acceptable quality and lowest cost — contains all of production: build the flow around the limiting step (the egg), fix problems at the lowest-value stage, batch where setup costs dominate, and choose deliberately between building to forecast and building to order. Every team — engineering, support, recruiting — runs a production line, whether or not anyone has drawn it.

Why it works: Knowledge work hides its assembly line, and invisible flow invites firefighting. Production thinking makes flow visible: once you know the limiting step, everything else gets scheduled around it; once defects are caught at the egg stage instead of on the customer's plate, fixing them costs a fraction.

Key insights:

  • Build around the limiting step: find the longest, hardest, or most expensive stage and offset everything else from it — often code review, staging access, or one overloaded specialist
  • Fix problems at the lowest-value stage: kill a flawed spec in review, not after three sprints of building on it
  • Batch work with high setup cost — interviews, code reviews, interrupt handling — so the setup amortizes across the batch
  • Most knowledge work is built to forecast, not to order: staff the pipeline to the forecast and accept controlled risk, as the toast goes down before the customer orders
  • You cannot watch all the work: treat it as a black box and cut windows into it with a handful of indicators

Applications:

ContextApplicationExample
Sprint flowSchedule around the limiting stepReview is the bottleneck → protect reviewer hours before starting new work
Quality gatesInspect at the lowest-value stageSpec review kills a flawed design before a three-week build
Hiring pipelineBatch and build to forecastPhone screens batched Tue/Thu; interviewer capacity staffed to the offer-date forecast

Ethical boundary: Run systems hot, not humans — production thinking optimizes the work, never treats people as interchangeable machines.

See: references/indicators-and-production.md when finding a limiting step or building a dashboard — limiting-step analysis, worked indicator pairs, leading vs trend indicators, stagger charts, and how to run an operation review.

2. Indicators That Don't Lie

Core concept: Measure output, not activity — what the team shipped that survived, not how busy it looked. Pair every quantity indicator with a quality counterpart so neither can be optimized at the other's expense, favor leading indicators that buy time to act, and report forecasts in stagger charts that show how each forecast evolved.

Why it works: People do what management measures, so an unpaired indicator is an instruction to game it. The pair closes the loop: push throughput and the escape rate exposes the corner-cutting. Leading indicators and stagger charts convert measurement from autopsy to steering.

Key insights:

  • Lines of code, hours logged, and tickets touched are activity; features alive in production and problems solved are output
  • Pair quantity with quality: deploys/week with change-failure rate, ticket closes with reopens, velocity with incident count
  • Leading indicators (review queue age, build flakiness, on-call page rate) warn before output drops; trend indicators compare output against your own history and forecast
  • A stagger chart re-forecasts the same horizon every period; reading down a column shows whether forecasting is honest, optimistic, or sandbagged
  • Administrative work measures like a factory: offers per recruiter-week, invoices processed per day — always with a quality pair

Applications:

ContextApplicationExample
Eng dashboardPair quantity with qualityDeploys/week paired with change-failure rate
Support opsOutput plus its quality shadowTickets resolved/day paired with reopen rate and CSAT
Quarterly forecastStagger chartRe-forecast quarter-end ARR monthly; drift visible down each column

Ethical boundary: Indicators measure the work, not the worker — used for surveillance, they teach people to optimize the number instead of the output.

3. Managerial Leverage

Core concept: Leverage is the output created per unit of managerial time. High-leverage activities affect many people at once (training, well-prepared decisions, information gathering) or redirect months of work with a small, well-timed nudge. The calendar is the manager's production system: forecast the key events, batch the rest, and say no at the source when capacity is full.

Why it works: Managerial activities differ by orders of magnitude in output per hour — ninety minutes preparing a review shapes a year of someone's work, while a day of meddling subtracts output. A manager who lets the calendar happen to them spends prime hours on whatever shouted loudest.

Key insights:

  • Negative leverage is real: meddling (supervising an expert in detail), waffling (stalling a decision others wait on), and a manager's visible gloom all multiply downward through the team
  • Delegate the tasks you know best — monitoring them costs you least — and remember that delegation without monitoring is abdication
  • Monitor at the task level, not the person level: sample like incoming inspection, deeper at low task-relevant maturity, lighter as it rises
  • Forecast your limiting steps: put 1:1s, staff meetings, reviews, and planning on the calendar first and let interrupts fill around them, not the reverse
  • Run below 100% load: a fully booked manager turns every surprise into a delay for everyone downstream; saying no early is cheaper than failing late
  • Batch interruptions with office hours and known checkpoints instead of letting them shred maker time

Applications:

ContextApplicationExample
Week designForecast fixed events, batch the rest1:1s Tue-Wed mornings, PR reviews batched daily at 4pm, Monday deep-work block
DelegationMonitoring depth by TRMNew hire's first migration: plan review plus daily spot checks; veteran's: rollout plan only
InterruptsConvert random pings to office hoursTwo daily drop-in slots replace ad hoc Slack escalations

Ethical boundary: Leverage means multiplying others' output, never hoarding information or approvals until you become the bottleneck everyone must visit.

See: references/leverage-and-calendar.md when auditing a calendar or setting up delegation — the weekly leverage audit, positive/negative-leverage catalog, delegation protocol with TRM-based monitoring depth, calendar-redesign procedure, and interruption management.

4. Meetings Are the Medium of Management

Core concept: A meeting is not a symptom of bad management; it is where managerial work — gathering information, imparting it, deciding, nudging — actually happens. Process-oriented meetings (one-on-ones, staff meetings, operation reviews) run on a regular cadence and should carry the bulk of that work, roughly a quarter of the calendar. Mission-oriented meetings are ad hoc and exist solely to produce a decision.

Why it works: Regularity makes meetings cheap — standing agendas, shared expectations, zero setup cost — and starves the expensive kind: issues get caught small in 1:1s and staff meetings instead of exploding into emergency decision meetings. Grove's malorganization test: ad hoc mission-oriented meetings eating more than about a quarter of managerial time means the process is broken.

Key insights:

  • The 1:1 is the subordinate's meeting: they own the agenda and bring it; the supervisor's job is to listen and learn what is really going on
  • Set 1:1 frequency by task-relevant maturity, not seniority or affection: new-to-task weekly, veterans every few weeks — never less than monthly
  • Both sides keep a "hold" list of non-urgent items for the next 1:1 — it batches interruptions away
  • The supervisor takes the notes: writing down agreed actions signals commitment and forces follow-up
  • "One more thing": after the agenda is done, ask what else is on their mind — the real issue often surfaces in the last five minutes
  • Staff meetings are controlled free discussion — the manager moderates as a Socratic prodder, not a lecturer; a recurring "ad hoc" meeting is a process meeting in denial

Applications:

ContextApplicationExample
New reportWeekly 1:1, their agendaFirst 90 days: 60 minutes weekly; agenda arrives the day before
Team syncControlled free discussionTwo-minute updates, then debate on two pre-flagged issues
Recurring "urgent" meetingConvert to processThird ad hoc incident review this month becomes a standing ops review

Ethical boundary: Hijacking the 1:1 for status extraction teaches people to stop bringing real problems — status belongs in writing.

See: references/meetings-and-one-on-ones.md when designing a meeting cadence or running a 1:1 — the full 1:1 playbook with agenda templates, staff-meeting design, operation-review roles, and meeting-cost math for when to kill a meeting.

5. Decisions and Planning (incl. OKRs)

Core concept: The ideal decision moves through free discussion (all views aired, dissent welcome), a clear decision (stated crisply — the more contentious, the crisper), and full support (disagree and commit). Decisions belong at the lowest competent level, closest to current technical knowledge. Planning runs the same arc: assess environmental demand, face present status honestly, close the gap — because the output of planning is decisions and actions taken now, not documents.

Why it works: Free discussion surfaces knowledge that lives at the edges; a clear decision prevents the costliest outcome, ambiguity; full support lets the organization move without unanimity. And today's firefight is yesterday's planning failure — planning works on next year's gap, not this week's smoke.

Key insights:

  • Peer-group syndrome — peers circling, waiting for someone senior to lean — is broken by peer-plus-one: one senior person in the room sanctioned to tip the decision
  • Before any decision meeting, answer six questions: what decision, by when, who decides, who is consulted, who ratifies or vetoes, who is informed
  • When no one person has both, pair the freshest technical knowledge with the strongest organizational judgment
  • Reversing a decision quietly is waffling; reversing it openly on new facts is management
  • MBO/OKRs answer two questions: where do I want to go (objective), and how will I pace myself to see I am getting there (key results)
  • Keep objectives few and key results measurable enough to score without argument; cascade so one level's key results become the next level's objectives — and never wire them mechanically to compensation

Applications:

ContextApplicationExample
Architecture choiceFree discussion → clear decision → commitRFC debated one week; tech lead decides; dissent recorded, then full support
Decision prepSix-question brief"Pick payments vendor by Jun 30; platform PM decides; eng and finance consulted; VP ratifies"
Quarterly planningCascading OKRsCompany KR "checkout p95 under 800ms" becomes the platform team's objective

See: references/decisions-planning-okrs.md when prepping a contentious decision or a planning cycle — the six-question brief, peer-group-syndrome counters, three-step planning, and a Grove-style OKR cascade with pitfalls.

6. Task-Relevant Maturity, Reviews, and Training

Core concept: There is no universally good management style. The right style depends on the subordinate's task-relevant maturity (TRM) — their experience, training, and confidence for this specific task: low TRM calls for structured "how" instruction, medium for mutual reasoning about "what and why", high for agreed objectives with light monitoring. TRM is task-specific, not seniority, so style must shift the moment the task does.

Why it works: Mismatched style fails in both directions — hands-off at low TRM is abandonment dressed as empowerment; detailed instruction at high TRM is meddling that destroys ownership. The performance review is where the cost of a mismatch compounds: a year's feedback delivered in the wrong register lands as either neglect or insult.

Key insights:

  • A star promoted into management is high-TRM on engineering and low-TRM on managing — structure the new task even for your best person
  • The performance review is the single most important form of task-relevant feedback a supervisor gives; its only purpose is improving the recipient's performance
  • Assess, don't blend: complete the written assessment first, then separately decide which three messages will actually change next year's output
  • No surprises: anything that startles the recipient in a review is the supervisor's failure, logged in public
  • The ace who is coasting deserves the most review effort — "keep it up" robs your best performer of their next level
  • Once lower needs are met, only an ever-rising, self-set bar motivates (the athlete mindset) — and training is the manager's highest-leverage way to raise that bar: deliver it yourself, because outsourcing training outsources standards

Applications:

ContextApplicationExample
Newly promoted managerRe-rate TRM per taskWeekly structured 1:1s on hiring and delegation, even for a star engineer
Review prepAssess first, message secondFull written assessment, then the three messages that change next year
Team capabilityManager-taught trainingEM personally teaches a four-session incident-response course

See: references/case-studies.md when preparing a review or coaching a newly promoted manager — three worked scenarios (meeting-drowned new manager, velocity-up/quality-down, a botched review repaired with TRM coaching).

Common Mistakes

MistakeWhy It FailsFix
Measuring activity, not outputBusyness is gameable and says nothing about resultsCount what shipped and survived; pair quantity with quality
Publishing unpaired indicatorsThe team optimizes the number at quality's expenseAdd the quality counterpart before the metric goes live
Skipping 1:1s when busyCancels the highest-leverage 90 minutes on the calendarTreat 1:1s as forecasted production steps: reschedule, never drop
Decisions by rankKnowledge lives at the lowest competent level; rank silences itFree discussion, then a clear decision by the named decider
OKRs as a compensation formulaGuarantees sandbagged, safe objectivesKeep OKRs a stretch tool; comp weighs more than OKR hit rate
One management style for everyoneAbandons the new, smothers the experiencedMatch style to task-relevant maturity, task by task
Catching defects at the highest-value stageCost multiplies at every stage a flaw survivesInspect specs and plans, not just production
Saving feedback for the annual reviewIt detonates all at once; trust and the year are both lostNo-surprises rule: deliver feedback when the event happens

Quick Diagnostic

QuestionIf NoAction
Can you state your team's output in one sentence?You are managing activityDefine output; build 4-6 indicators around it
Does every quantity metric have a quality pair?The number is being gamed alreadyPair it: throughput with escapes, closes with reopens
Do you know your team's limiting step?Flow is built around the wrong constraintFind where work queues longest; schedule around it
Did your reports set the agendas of their last 1:1s?You ran status meetings insteadHand the agenda to the subordinate; you take the notes
Is 1:1 frequency set by task-relevant maturity?Someone is over- or under-managedWeekly for new-to-task, monthly for veterans
Was your last big decision made at the lowest competent level?Rank decided; knowledge watchedName decider, consulted, and ratifier before the meeting
Would your team set the same OKRs if pay weren't attached?Objectives are sandbaggedDecouple OKRs from the compensation formula
Have you personally taught your team anything this quarter?Highest-leverage activity skippedSchedule a manager-taught course now

Further Reading

About the Author

Andrew S. Grove (1936-2016) fled Hungary at twenty, became Intel's third employee, and rose to president, CEO, and chairman, driving the company's famous pivot from memory chips to microprocessors. Time's 1997 Man of the Year, he mentored a generation of Silicon Valley leaders, and his management-by-objectives system became the OKR method now standard across tech.

附带文件

references/case-studies.md
# Case Studies: High Output Management in Practice

## Table of Contents

- [Case Study 1: The Manager Drowning in Meetings](#case-study-1-the-manager-drowning-in-meetings)
- [Case Study 2: Velocity Up, Quality Down](#case-study-2-velocity-up-quality-down)
- [Case Study 3: The Botched Performance Review](#case-study-3-the-botched-performance-review)
- [Key Takeaways](#key-takeaways)

## Case Study 1: The Manager Drowning in Meetings

### Context

Dana, a strong backend engineer at a 120-person fintech, was promoted to engineering manager of a seven-person team. Eight months in, she was working 55-hour weeks: 34 hours of meetings, the rest fragmented into slivers she spent reviewing PRs at night "to stay technical." Her team's delivery had slowed, two engineers were quietly job-hunting, and her own manager labeled the team "a black box."

### The Problems

**No process meetings, all ad hoc.** Dana held no regular 1:1s ("no time") and no staff meeting. Information reached her through interruptions — 40+ Slack pings a day — and through emergency meetings that existed because problems were caught late. The absence of process meetings was *creating* the ad hoc ones.

**Doing instead of managing.** A week log showed 14 hours of IC work: PR reviews she grabbed first, two production fixes she did herself because "it's faster," and a vendor integration she hadn't handed off. Meanwhile, delegation-shaped work — a design doc her senior engineer could own, interview loops, the on-call rotation redesign — sat with her.

**Negative leverage, invisible.** Three decisions (a schema change, a library upgrade, a hire) had been waiting on her for over two weeks, blocking four people. Her habit of "jumping in to help" on tasks her senior engineers owned was reread by them as distrust.

### The Intervention

**Week 1: Log and classify.** Dana logged five days at 30-minute granularity, then classified each block (information gathering / giving, decision-making, nudging, role model, *doing*) and scored leverage. Results: 11% high-leverage, 31% doing, 12% negative leverage (stalled decisions, meddling, unprepared meetings), the rest neutral. The three stalled decisions alone were blocking an estimated 30 engineer-days.

**Week 2: Install the process meetings.** Weekly 60-minute 1:1s with all seven reports (everyone was effectively new to *her*, and two were new to their tasks), each with a subordinate-owned agenda template and a hold list; one weekly 75-minute staff meeting with a metrics minute, a round, and two debated topics. She cleared the three stalled decisions in the first staff meeting using the six-question brief — two she delegated outright with a named decider.

**Weeks 2-3: Redesign the calendar as a production system.** Fixed events first (1:1s Tue/Wed mornings, staff meeting Monday after lunch); PR review batched to one 4pm window with a rule that she reviews only designs and risky changes, not routine PRs; two maker blocks; office hours 2-3pm daily replacing always-on Slack; 15% left unscheduled. Standing meetings she merely attended got the cost test — she exited four of them, sending a written update instead.

**Weeks 3-8: Delegate with TRM-based monitoring.** The vendor integration went to her senior engineer (high TRM: monitoring = rollout plan review only). The on-call redesign went to a mid-level engineer (medium TRM: weekly check on the plan plus spot checks). She wrote down the monitoring plan in each handoff conversation, and announced — publicly — that she would stop reviewing routine PRs.

### Results After 8 Weeks

| Metric | Before | After |
|--------|--------|-------|
| Meeting hours/week (Dana) | 34 | 19 (9 process, 10 other) |
| Ad hoc "urgent" meetings/week | 6-8 | 1-2 |
| Slack interruptions/day | 40+ | ~12 (office hours + hold lists) |
| Decisions pending >1 week | 3 | 0 |
| Dana's IC "doing" hours | 14 | 4 (design reviews only) |
| Team features shipped/sprint | 2-3 | 4-5 |
| Hours/week (Dana) | 55 | 44 |

The two job-hunting engineers stayed; both later said the 1:1s were the reason — problems they had assumed she didn't care about turned out to be hold-list items she now acted on.

### Lessons Learned

1. **The absence of process meetings creates the meeting overload.** The ad hoc meetings were the symptom; installing 1:1s and a staff meeting removed their cause.
2. **The log doesn't lie.** Dana guessed she spent "a few hours" on IC work; it was 14. Leverage cannot be improved before it is measured.
3. **Stalled decisions are the most expensive line item.** Twelve percent of her week was negative leverage, and most of its cost landed on other people's calendars.
4. **Delegation needed a published monitoring plan** — once monitoring was announced as task-level QA rather than improvised check-ins, her seniors stopped reading it as distrust.

## Case Study 2: Velocity Up, Quality Down

### Context

A nine-person product team at a B2B SaaS company adopted velocity (story points per sprint) as its headline metric after a slow quarter. Leadership praised rising numbers in all-hands. Two quarters later velocity was up 40% — and the team was miserable: incidents up, on-call brutal, and a key customer threatening to churn over reliability.

### The Problems

**A single unpaired indicator.** Velocity was quantity with no quality counterpart. Points rose exactly the way unpaired numbers always rise: tests skipped, migrations deferred, reviews rubber-stamped, stories inflated. Nobody was cheating consciously; the team was doing what management measured.

**The damage was visible only in unmeasured places.** Change-failure rate had doubled and sev-2 incidents went from two to five per month — but neither was on the dashboard, so velocity reviews stayed celebratory while on-call quietly burned out.

**The limiting step was being flooded, not fixed.** Code review was the constraint (median 31 hours to first review). Pushing more stories into the sprint didn't raise output; it raised WIP, pressure to rubber-stamp, and escapes.

**Forecasts had become theater.** Sprint commitments were set to impress and missed by 20-30%, then quietly re-explained. No record of forecast vs actual existed.

### The Intervention

**Step 1: Find the limiting step.** The team mapped its flow (spec → build → review → QA → deploy) and measured queue times for the last 30 stories. Review queues dominated: work waited 31 hours median, with WIP piling in front. Responses: reviewer hours protected before new work starts, a WIP limit upstream of review, linters and CI taking mechanical findings off reviewers, and risky-change review batched into a daily window.

**Step 2: Build the paired dashboard.** Four indicators, shown only together: features shipped per sprint (quantity) with change-failure rate (quality pair); time-to-first-review (quantity) with defect escape rate per 100 merges (quality pair). Two leading indicators alongside: review queue age and on-call pages per week, each with a pre-committed action threshold.

**Step 3: Stagger-chart the forecasts.** Each sprint, the team re-forecast the next three sprints' completed scope, keeping every prior forecast visible. The first month exposed a systematic 25% over-forecast — discussed openly, without blame, as a bias to correct.

**Step 4: Monthly operation review.** The lead presented the paired trends to the wider org — reviewing manager briefed to praise honest misses. The first review opened with the lead stating the quality cost of the velocity push before anyone asked.

### Results After Two Quarters

| Metric | Peak "velocity era" | After |
|--------|---------------------|-------|
| Story points/sprint | 58 | 49 |
| Change-failure rate | 14% | 5% |
| Sev-2 incidents/month | 5 | 1-2 |
| Median time-to-first-review | 31 h | 7 h |
| Defect escapes per 100 merges | 9 | 3 |
| Sprint forecast error | -25% (over-forecast) | -6% |
| On-call pages/week | 19 | 6 |

Velocity dropped 15% and nobody minded: features alive and stable in production — the team's actual output — rose, and the churn-threatening customer renewed.

### Lessons Learned

1. **Any indicator pushed hard will be achieved; the only question is what is sacrificed.** The pair makes the sacrifice visible before customers report it.
2. **The limiting step, not effort, sets output.** Flooding a constrained flow with more work converts effort into queues and defects.
3. **Stagger charts turn forecast bias into a measured, fixable quantity** — and honesty about forecasts proved contagious into estimates, reviews, and postmortems.
4. **Operation reviews changed incentives upward**: once leadership saw paired trends, "velocity up" stopped being praiseworthy on its own — the metric's audience, not the team, had been the root incentive problem.

## Case Study 3: The Botched Performance Review

### Context

Marcus, a staff engineer and the acknowledged ace of a data platform team, received a "meets expectations minus" annual review from his manager, Lena. He was blindsided — every prior signal had been positive — and furious. He stopped volunteering in design reviews, and his calendar began showing recruiter calls. Lena's own manager asked her to repair it.

### The Problems

**Total surprise.** Lena had been dissatisfied for months: Marcus's new charter (leading the streaming-platform migration — coordination, mentoring, stakeholder work) was going badly, while he kept retreating to the batch-pipeline work he was brilliant at. She had said nothing in their sporadic 1:1s. The review was the first time he heard any of it — a supervisor's failure, logged in public.

**Blended message.** The written review mixed praise and criticism so thoroughly ("exceptional technical depth, though stakeholders sometimes…") that drafting it had felt safe — and reading it felt incoherent. The rating contradicted the prose. Marcus latched onto the praise, concluded the rating was arbitrary, and assigned it to politics.

**A TRM misread at the root.** Lena had reasoned: staff engineer, ten years' experience, needs no support. But TRM is task-specific. On batch pipelines Marcus's TRM was the highest in the org; on cross-team migration leadership it was *low* — he had never done it. Lena's hands-off style was right for the old task and was abandonment on the new one. He had been coasting on the ace task partly because no one had structured the one he was failing at — the classic ace-who-coasts pattern, mishandled.

**Comp drove the message.** The "minus" existed mostly to justify a budget-constrained comp outcome, inverting the review's purpose: assessment had been reverse-engineered from money rather than performance.

### The Repair

**Step 1: Assess, don't blend — in writing, first.** Lena rewrote the assessment before scheduling any meeting: outcomes only, in two explicitly separated lanes. Batch platform: exceptional, with named results. Migration leadership: not delivering — milestones missed, two partner teams escalating, mentoring not happening — with named instances. Then she chose the three messages that would change next year's output: (1) the migration is now the job and it is going badly; (2) the cause is missing skills for a new kind of task, not effort or talent; (3) here is the structure we'll build, because the goal is for you to lead at the next level.

**Step 2: Own the failure, then deliver straight.** In the repair conversation Lena opened with her own miss: "You learned this in a review instead of in March. That was my failure, and the no-surprises rule starts now." Then the three messages, undiluted — no praise sandwich. Marcus argued; Lena listened fully (the heart-to-heart a review requires), and did not retract the assessment.

**Step 3: TRM-matched structure.** Together they re-scoped the work explicitly: migration leadership treated as a low-TRM task — weekly structured 1:1s on it (stakeholder map, milestone plan, a "what/when/how" level of detail that would have insulted him on pipeline work and was relief here), Lena attending his first partner-team negotiations as observer-coach, and a two-session course Lena taught herself on running cross-team programs. Batch-pipeline work stayed high-TRM: objectives and monitoring only.

**Step 4: No surprises, ever again.** Task-relevant feedback moved to the moment of the event — wins and misses named in the week they happened, logged in the shared 1:1 doc. A six-month interim review was scheduled in writing, with the explicit promise that nothing in it would be new.

### Results

| Signal | At the botched review | Six months later |
|--------|----------------------|------------------|
| Migration milestones | 2 of 5 hit | 5 of 5 hit |
| Partner-team escalations | 2 open | 0 |
| Marcus's 1:1 cadence on migration work | Sporadic | Weekly, his agenda |
| Engineers Marcus is mentoring | 0 | 2 |
| Interim review surprises | — | None, by design |
| Retention | Recruiter calls | Took the senior-staff track; stayed |

The interim review rated the migration work "exceeds" — and contained, verbatim, sentences Marcus had already heard in 1:1s. He later told Lena the original review was the most useful bad day of his career, "but only because of what came after it."

### Lessons Learned

1. **A surprise in a review is always the supervisor's failure.** The review is where accumulated, already-delivered feedback is consolidated — never where it debuts.
2. **Assess, don't blend.** Write the full assessment first, then choose the few messages that change next year's output. Blending to make delivery comfortable makes the message incoherent.
3. **TRM is task-specific, and promotions reset it.** The org's best engineer was a beginner at the new task; structured management there was coaching, not condescension.
4. **Your ace deserves the most review effort, not the least.** "Keep it up" robs the top performer of their next level — and the coasting ace is usually a structure problem before it is an attitude problem.

## Key Takeaways

**1. Output is the only judge.** Every intervention above was scored the same way: did the organization's output rise? Meetings, indicators, reviews, and OKRs are machinery for that, not deliverables in themselves.

**2. Measure before moralizing.** Dana's week log, the team's queue times, and Lena's written assessment all replaced a story ("I'm just busy", "we're faster", "he's difficult") with data — and the data redirected the fix every time.

**3. Pair every number and forecast in the open.** Unpaired indicators manufactured the velocity crisis; stagger charts and paired dashboards cured it. Honesty is a system property before it is a virtue.

**4. Style follows task-relevant maturity.** The same person needs structure on one task and autonomy on another, simultaneously. Most "people problems" in these cases were style-to-TRM mismatches wearing personality costumes.

**5. Process meetings are where problems get caught small.** 1:1s with hold lists, staff meetings with debated topics, and operation reviews with honest trends starved the emergency meetings and the year-end detonations alike.
references/decisions-planning-okrs.md
# Decisions, Planning, and OKRs

## Table of Contents

- [The Ideal Decision Process](#the-ideal-decision-process)
- [The Decision Brief: Six Questions](#the-decision-brief-six-questions)
- [The Lowest Competent Level](#the-lowest-competent-level)
- [Countering Peer-Group Syndrome](#countering-peer-group-syndrome)
- [The Three-Step Planning Process](#the-three-step-planning-process)
- [Writing OKRs Grove-Style](#writing-okrs-grove-style)
- [A Worked OKR Cascade](#a-worked-okr-cascade)
- [Pitfalls](#pitfalls)

## The Ideal Decision Process

Grove's model has three stages, in strict order, with no skipping:

**1. Free discussion.** Every relevant view and fact gets on the table, including — especially — dissent. "Free" is literal: people argue the merits regardless of rank, half-formed worries are admissible, and the senior person's job is to *draw out* disagreement, not to win early. Two disciplines make it real: senior people speak last (the moment the boss leans, the discussion is over whether anyone admits it), and dissent is framed as an obligation, not a courtesy — staying silent in the discussion and critical in the hallway is the one unforgivable move.

**2. Clear decision.** Discussion converges or time runs out; either way, the named decision-maker states the decision crisply, in writing, with the reasoning. Grove's observation: the more contentious the issue, the *more* unambiguous the wording must be, precisely because everyone is motivated to hear their preferred version. Vague decisions are the most expensive output a meeting can produce — everyone leaves to execute a different one.

**3. Full support.** Everyone commits to the decision's success — *support*, not pretended agreement. "Disagree and commit" is honorable on both sides: the dissenter argued fully and now executes fully; the organization in turn owes them that the disagreement was genuinely heard. Relitigating in side channels is sabotage; reopening *openly* because material new facts arrived is management. The difference is the venue and the honesty.

The three stages also define the failure modes: discussion without decision (the circling committee), decision without discussion (rank ruling on partial knowledge), and decision without support (the quiet veto in execution).

## The Decision Brief: Six Questions

Grove insisted a decision be framed *before* the meeting that makes it. Answer six questions in writing and attach them to the invite:

```
DECISION BRIEF

1. What decision is needed?
   One sentence, phrased as a question with identifiable options.
2. By when?
   The real deadline and what it's anchored to (contract, launch, hiring window).
3. Who decides?
   One name. Not a committee.
4. Who must be consulted before the decision?
   The people whose knowledge or stake earns them input — with a date for it.
5. Who ratifies or can veto?
   Usually the decider's manager; ratification checks blast radius, not taste.
6. Who must be informed after?
   Everyone whose work changes because of the outcome.
```

**Worked example:**

```
1. Decision: Which payments provider for EU expansion — extend current
   vendor, or migrate to Adyen?
2. By when: June 30 (contract renewal is July 15; migration lead time 6 weeks).
3. Who decides: Priya (platform PM).
4. Consulted: payments eng lead (integration cost), finance (fees model),
   support lead (dispute tooling) — input by June 20.
5. Ratifies: VP Engineering (commits 2 engineers for a quarter if migrating).
6. Informed: checkout team, data team (settlement pipelines), exec staff.
```

The brief takes fifteen minutes to write and routinely saves weeks: half the time, writing it reveals the meeting is unnecessary (the decider can decide today), mis-staffed (the real consultations haven't happened), or premature (the deadline is invented).

## The Lowest Competent Level

Decisions should be made at the lowest level where someone has both the relevant technical grasp and acceptable judgment about consequences. Two reasons. First, knowledge currency: in fast-moving fields the people closest to the work hold the freshest technical truth; every level of escalation trades current knowledge for older generalizations. Second, speed and ownership: decisions made by the people who must execute them start with built-in commitment.

When no single person has both halves, **compose the decision**: pair the engineer with the freshest technical knowledge and the manager carrying organizational judgment, and have them decide *together* — explicitly, as named co-deciders, not as "input" flowing upward to be overruled. Grove considered this blend of knowledge-power and position-power the everyday business of a well-run company.

What escalates: decisions whose blast radius genuinely exceeds the local view (cross-org resource shifts, irreversible commitments, precedent-setting calls). What does not: decisions escalated because someone senior *would like* to make them. Each unnecessary escalation teaches the team that authority, not knowledge, decides — and the best people start pre-clearing everything, which is how organizations get slow.

## Countering Peer-Group Syndrome

Put six peers in a room with a contested question and watch: everyone hedges, nobody wants to stick out with a strong position that might lose, and the discussion circles — not because nobody knows, but because nobody wants to be *wrong in front of equals*. Grove named it peer-group syndrome, and its root is fear of looking dumb, which silences exactly the people with the most current knowledge.

Counters, in order of power:

- **Peer-plus-one.** Add one person senior to the group, sanctioned in advance to break ties and absorb the risk of the call. Their presence licenses strong positions: someone in the room can bless one.
- **The chairman states the question first.** Circling thrives on ambiguity about what is being decided. Opening with the decision brief's question collapses the fog.
- **Written positions before the meeting.** One paragraph from each participant, circulated with the pre-read. Positions taken in writing before social pressure exists are more honest, and the spread of views is visible immediately.
- **Senior people speak last** — and ask questions before stating views.
- **Assign the counter-case.** Name someone to argue the strongest version of the losing side, decoupling the argument from the arguer's reputation.
- **Reward visible mind-changing.** When someone updates on evidence, the senior person marks it as strength. Once changing your mind is safe, taking a position is too.

## The Three-Step Planning Process

Grove's planning frame, applicable to a yearly org plan or a quarter's team plan:

**Step 1 — Environmental demand: what will be wanted of you?** Not what you want to build — what your environment (customers, adjacent teams who depend on you, the market, the company's direction) will demand of you over the planning horizon, typically the next one to two years. List the demands and their trajectory: growing, flat, fading. The discipline is *outside-in*: a platform team's environment is its internal customers' roadmaps; a product team's is the market and the support queue.

**Step 2 — Present status: what are you producing now?** Current capabilities and trajectory, stated with uncomfortable honesty: what is actually shipping, what is in flight *and will genuinely finish*, what is in flight and will not (say so now, not in month eleven), and where capacity is actually going (run the leverage audit's numbers — firefighting and maintenance load included).

**Step 3 — Close the gap.** The difference between step 1 and step 2 is the gap, and the plan is the set of *actions taken now* — projects started, projects killed, hires opened, skills built — to close it. Two Grove rules govern this step:

- **The output of planning is decisions and actions, not documents.** A planning process that produces a deck and no changed behavior produced nothing. The test of a finished plan: what did we *start*, *stop*, and *commit to* this week because of it?
- **Today's gap reflects yesterday's planning failure.** If you are firefighting now, the fire was set by what last year's plan missed. Therefore plan for the *next* gap: today's actions affect output one to two years out. A plan addressed to this week's smoke is reaction wearing planning's clothes.

**Worked sketch:** a platform team's environment scan shows three product teams shipping mobile features next year (demand: mobile-ready APIs, rising) and the company entering the EU (demand: data residency, new and hard). Present status: the API gateway is web-centric; one engineer understands the data layer; 40% of capacity goes to toil from a legacy queue system. Gap-closing actions decided now: kill the legacy queue this quarter (frees the 40%), open one hire with data-residency experience, start the gateway redesign in Q3 — and explicitly *not* pursue the internal analytics tool two teams asked for, with the refusal communicated and dated. That last item matters: a plan that declines nothing has decided nothing.

## Writing OKRs Grove-Style

Grove's management by objectives — which John Doerr carried from Intel to Google as OKRs — answers exactly two questions:

1. **Where do I want to go?** The *objective*: directional, motivating, few in number.
2. **How will I pace myself to see if I am getting there?** The *key results*: measurable milestones, stated so you can answer "did I hit it?" with yes or no, **without argument**.

Grove's rules of construction:

- **Few objectives.** Two or three per level per cycle... and that is not a typo. Each objective you add steals attention from the others; an OKR set with seven objectives is a to-do list with ambitions.
- **Key results are verifiable, dated milestones**, not activities. "Engage with the migration project" is not a key result; "legacy queue serving 0% of production traffic by Nov 15" is.
- **Short cadence.** Quarterly objectives with a monthly look — the system exists to provide *feedback during the race*, not a grade after it.
- **Cascaded, not dictated.** One level's key results supply the next level's candidate objectives — but each team *writes its own*, and a healthy share of objectives flow bottom-up from the people closest to the work. Alignment comes from visibility and negotiation, not transcription.
- **A stretch instrument, not a contract.** Missing a genuinely ambitious key result while performing well is a fine outcome; hitting 100% of everything means the bar was set for safety. The review of a person weighs far more than their OKR scorecard — which is why OKRs must never be wired mechanically to compensation.

## A Worked OKR Cascade

**Company (quarter):**
- **Objective:** Make checkout the fastest in the mid-market segment.
  - KR1: p95 checkout latency under 800ms (from 2.1s).
  - KR2: Checkout conversion +2 points.
  - KR3: Zero sev-1 incidents in the checkout path this quarter.

**Platform team** (adopts KR1 as its objective — the cascade joint):
- **Objective:** Cut checkout p95 below 800ms.
  - KR1: Payment-provider calls parallelized; sequential wait eliminated by Aug 15.
  - KR2: Edge-cached session auth live in EU and US regions by Sep 1.
  - KR3: Latency budget dashboard with per-service attribution adopted by all three checkout services by Jul 31.

**Individual engineer** (adopts team KR1 as her objective):
- **Objective:** Eliminate sequential payment-call latency.
  - KR1: Parallel orchestration design ratified in RFC review by Jul 10.
  - KR2: Shipped behind a flag, 10% traffic, error rate within 0.1% of baseline by Aug 1.
  - KR3: 100% rollout with p95 contribution under 300ms by Aug 15.

Each level is written by its owner, each key result is a dated yes/no, and reading upward, any engineer can trace why her work matters to the company's quarter — Grove's definition of the system working.

## Pitfalls

| Pitfall | What it looks like | Correction |
|---------|--------------------|------------|
| OKRs wired to compensation | Objectives sandbagged to safely-hittable; stretch vanishes | Comp reviews weigh whole performance; OKR scores are one input at most, never a formula |
| 100% attainment celebrated | The bar was set where it couldn't be missed | Treat ~70-80% on genuine stretch as healthy; investigate perfect quarters like misses |
| Objective inflation | Six-plus objectives per team | Cap at three; the cut list is the strategy |
| Activity key results | "Work on", "support", "continue" | Rewrite as dated, verifiable outcomes |
| Cascade as dictation | Teams transcribe their slice from above | Each level writes its own; expect bottom-up objectives too |
| Set-and-forget | OKRs written in week 1, reread in week 13 | Monthly check against the stagger chart; re-forecast, don't re-write history |
| Decision relitigated in hallways | "Supported" decision quietly starved in execution | Name it: reopen openly with new facts, or commit — there is no third venue |
| Planning produces a binder | Beautiful deck, unchanged behavior | End planning with started/stopped/committed actions, owners, and dates |
references/indicators-and-production.md
# Indicators and Production Principles for Software Teams

## Table of Contents

- [The Breakfast Factory, Translated](#the-breakfast-factory-translated)
- [Finding the Limiting Step](#finding-the-limiting-step)
- [Inspect at the Lowest-Value Stage](#inspect-at-the-lowest-value-stage)
- [Choosing Team Indicators](#choosing-team-indicators)
- [Pairing Indicators: Three Worked Examples](#pairing-indicators-three-worked-examples)
- [Leading vs Trend Indicators](#leading-vs-trend-indicators)
- [The Stagger Chart](#the-stagger-chart)
- [Running an Operation Review](#running-an-operation-review)
- [Measuring Administrative Work](#measuring-administrative-work)
- [Anti-Gaming Rules](#anti-gaming-rules)

## The Breakfast Factory, Translated

Grove opens the book with a waiter's problem: deliver a three-minute egg, buttered toast, and hot coffee to the table simultaneously, fresh, at acceptable cost. Everything in production is in that sentence — and everything in running a software team is too:

| Breakfast factory | Software team |
|-------------------|---------------|
| The egg (longest, least flexible step) | The limiting step: code review queue, staging environment, the one person who knows the billing system |
| Toast must start before the order is certain | Build to forecast: hiring, capacity, and roadmaps start on predicted demand |
| Candle the eggs before boiling | Inspect specs and designs before the build, not after |
| Batch toast in the toaster's capacity | Batch reviews, interviews, deploys where setup cost dominates |
| The waiter can't watch the kitchen continuously | Black-box the work; cut windows into it with indicators |

The discipline is to draw the team's actual production line — idea → spec → build → review → test → deploy → operate, or ticket → triage → fix → verify → close — and then manage the flow, not the individual heroics inside it.

## Finding the Limiting Step

The limiting step is the stage that is longest, most expensive, or hardest to scale. The whole flow should be built around it, because output equals the limiting step's throughput no matter how fast everything else runs.

**Procedure:**

1. **Draw the stages** of one unit of work from request to "alive in production." Six to eight boxes is the right altitude.
2. **Measure queue time at each boundary** for the last 20-30 units of work — time *waiting* between stages, not time being worked. Git timestamps, ticket histories, and PR metadata usually contain all of it.
3. **The stage with the longest queue in front of it is the limiting step.** Work piles up in front of constraints; that is the whole diagnostic.
4. **Verify with a thought experiment:** if this stage doubled its throughput, would the team's output rise? If yes, it is the constraint. If output would just pile up at the next stage, keep looking.

**Then build around it:**

- **Protect its capacity.** If review is limiting, reviewer hours are scheduled before new feature work, not squeezed after.
- **Stop starting work the constraint cannot absorb.** A WIP limit upstream of the limiting step is the software equivalent of not cracking eggs you cannot boil.
- **Offload it.** Move work off the constraint: linters and CI take mechanical findings off reviewers; templates take routine answers off the senior engineer.
- **Re-find it quarterly.** Constraints move once relieved. Teams that "fixed review" in Q1 are often staging-limited by Q3.

## Inspect at the Lowest-Value Stage

A flaw costs more at every stage it survives. The rule: detect and fix any problem at the lowest-value stage possible.

| Stage caught | Typical cost to fix |
|--------------|---------------------|
| One-page spec review | An hour and a conversation |
| Design/RFC review | A day and a revision |
| PR review | Days — code exists, sunk cost argues back |
| QA / staging | A week — context reload, retest |
| Production incident | Weeks — plus customers, trust, and on-call burnout |

Three kinds of inspection, mapped from the factory:

- **Incoming inspection:** requirements and specs. Is the problem real, the approach sound, the scope bounded? The cheapest hour in engineering is the hour spent killing a bad spec.
- **In-process inspection:** design reviews, PR review, CI. Catch flaws while the material is still cheap to rework.
- **Final inspection:** release checklists, canary deploys, smoke tests. Necessary, but if final inspection is your primary quality mechanism, you have built a factory that ships rotten eggs to the plating station.

Choose between **gate** (work stops until it passes — right for irreversible or high-blast-radius changes: schema migrations, auth, billing) and **monitoring** (work proceeds, samples are checked, drift triggers tightening — right for everything else, because gates everywhere destroy flow). A useful default: gates at incoming and final, monitoring in process — tightened temporarily where escapes have actually occurred, then loosened again. A variable-inspection scheme beats permanent maximum inspection.

## Choosing Team Indicators

Indicators are the windows cut into the black box. Choosing them well:

1. **Four to six, no more.** Beyond that, attention diffuses and nobody steers.
2. **Each measures output or a direct precondition of output** — not effort, not motion.
3. **Each has an owner and a review cadence** (the weekly team review and the monthly operation review below).
4. **Each quantity indicator gets a quality pair** — this is non-negotiable, because people will do what management measures, and an unpaired number is an instruction to game it.
5. **Cheap to collect.** An indicator that takes an afternoon to assemble dies in a month. Pull from systems (git, CI, ticketing, observability), not from human memory.

## Pairing Indicators: Three Worked Examples

**1. Code review throughput vs defect escape rate.**
A platform team measured time-to-first-review (median 26 hours) and made it the headline metric. Within six weeks the median fell to 4 hours — and rubber-stamp approvals rose with it; "LGTM" reviews with zero comments went from 18% to 55%, and bugs found after merge climbed. The fix was the pair: *PRs reviewed per week* and *time-to-first-review* displayed only alongside *defect escape rate* (bugs attributed to merged PRs per 100 merges, found within 30 days). Reviewers could no longer win by waving work through; the pair forced the real goal — fast *and* sound review. Stabilized at 8-hour first response with escapes at half the original rate.

**2. Ticket closes vs reopen rate.**
A support team paid attention (and a spiff) to tickets closed per agent-day. Closes rose 30%; so did "resolved" tickets that customers immediately reopened — agents were closing on first response with a boilerplate answer. Pairing *closes per agent-day* with *reopen rate within 7 days* and *CSAT on closed tickets* exposed the pattern in the first week: the two agents with the highest close counts had the worst reopen rates. Coaching, not punishment, followed — and the indicator pair was published to the team so everyone could see that the game had changed from "close fast" to "close once."

**3. Feature velocity vs incident rate.**
A product team celebrated rising velocity (story points per sprint, up 40% over two quarters). The same period: change-failure rate doubled, two sev-2 incidents per month became five, and on-call escalations rose. Velocity was being purchased with skipped tests and deferred migrations — invisible because nobody graphed the pair. The fix: a four-indicator dashboard — *features shipped per sprint*, *change-failure rate*, *incidents attributed to recent changes*, *p95 cycle time* — reviewed together in the monthly operation review. Velocity dropped 15% the next quarter; incidents fell by two-thirds; net output (features alive and stable in production) rose.

The general law: any indicator pushed hard will be achieved — the only question is what gets sacrificed to achieve it. The pair makes the sacrifice visible before the customer reports it.

## Leading vs Trend Indicators

**Leading indicators** let you act before output falls. Good ones for software teams: review queue age, build flakiness rate, on-call pages per week, backlog age of sev-2 bugs, recruiting pipeline depth, and a simple morale pulse. Each needs a believed threshold — a level at which you have pre-committed to act, otherwise you will explain away the warning ("the linearity index dipped, but surely it'll catch up").

**Trend indicators** show output against two baselines: your own history (deploys this month vs the last six) and your forecast (what you said you would do). Measurement against forecast is what turns an indicator from a mood into a commitment — which is the stagger chart's job.

## The Stagger Chart

A stagger chart re-forecasts the same horizon every period, keeping every previous forecast visible. Forecast the next three sprints' completed scope (or the quarter's ARR, or the month's hiring), and each period add a new row:

| Forecast made | Sprint 10 | Sprint 11 | Sprint 12 | Sprint 13 |
|---------------|-----------|-----------|-----------|-----------|
| In sprint 9 | 24 pts | 26 | 27 | — |
| In sprint 10 | **21 (actual)** | 24 | 26 | 27 |
| In sprint 11 | | **20 (actual)** | 23 | 26 |
| In sprint 12 | | | **22 (actual)** | 24 |

Read **down a column**: sprint 12 was forecast at 27, then 26, then 23, landing at 22. The team systematically over-forecasts by ~20% — visible in one glance, unarguable, and correctable. A team that sandbaggs shows the opposite signature (forecasts rising to meet comfortable actuals). The stagger chart does not punish misses; it makes forecast *bias* a measured, improvable quantity — which is the honest foundation under every commitment the team makes outward.

## Running an Operation Review

The operation review is where indicators meet an audience: managers present their area to people who do not see their work day-to-day — adjacent teams, skip-levels, new hires. Grove's purposes: it teaches, it motivates (people raise their game when their work has an audience), and it lets seniors sanity-check trends juniors might rationalize.

**Cast:** an *organizing manager* (logistics, agenda, timekeeping), a *reviewing manager* (senior; asks questions, sets the tone, never ambushes), 2-4 *presenters* (line managers/leads with their indicators), and the *audience* — whose job is to engage, not spectate.

**Cadence:** monthly for a department, quarterly for an org. Sixty to ninety minutes.

**Agenda template:**

1. (5 min) Reviewing manager: why we are here, what changed since last time.
2. (15 min × 3) Each presenter: their 4-6 indicators as *trends with forecasts* (stagger charts, not snapshots), one problem they are working, one ask of the room.
3. (10 min) Open questions across areas — the cross-pollination is the point.
4. (5 min) Reviewing manager: themes, decisions taken, actions with owners.

**Presenter rules:** trends, not points-in-time; pairs shown together; misses stated before anyone asks; no slide with more than one chart. **Reviewing-manager rules:** ask the second question ("what's underneath that?"), praise honest bad news, and never let the room punish a candid miss — one ambushed presenter converts the whole org's reviews into theater permanently.

## Measuring Administrative Work

Grove insisted administrative and knowledge work be measured like production, because it is production:

| Function | Output indicator | Quality pair |
|----------|------------------|--------------|
| Recruiting | Offers extended per week | Offer-accept rate; 90-day retention of hires |
| Support | Tickets resolved per agent-day | Reopen rate; CSAT |
| Documentation | Docs shipped/updated per month | Search success rate; support tickets on documented topics |
| Finance ops | Invoices processed per day | Error/dispute rate |
| IT/internal tools | Requests fulfilled per week | Repeat-request rate; requester satisfaction |

Same rules as engineering indicators: output not activity, paired, owned, trended against forecast.

## Anti-Gaming Rules

- **Publish the pair or publish nothing.** A quantity indicator released alone is a gaming instruction with a dashboard.
- **Indicators describe the work, not the worker.** Use team-level indicators for steering; individual performance is assessed in reviews with full context, not read off a throughput chart. The moment indicators become surveillance, people optimize the number and hide the truth — and you lose both the output and the warning system.
- **Never convert an indicator directly into compensation.** The instant money attaches, the indicator stops measuring reality (see the OKR pitfalls in [decisions-planning-okrs.md](decisions-planning-okrs.md)).
- **Expect Goodhart drift and rotate emphasis.** Any measure pushed for quarters degrades; re-derive indicators from the current limiting step, not from habit.
- **Let the team see everything you see.** Indicators reviewed in the open steer; indicators reviewed privately breed paranoia and creative accounting.
references/leverage-and-calendar.md
# Managerial Leverage and the Calendar

## Table of Contents

- [The Output Equation](#the-output-equation)
- [Step 1: Log a Real Week](#step-1-log-a-real-week)
- [Step 2: Classify Every Block](#step-2-classify-every-block)
- [Step 3: Score the Leverage](#step-3-score-the-leverage)
- [The Leverage Catalog](#the-leverage-catalog)
- [The Delegation Protocol](#the-delegation-protocol)
- [Monitoring Depth by Task-Relevant Maturity](#monitoring-depth-by-task-relevant-maturity)
- [Redesigning the Calendar](#redesigning-the-calendar)
- [Managing Interruptions](#managing-interruptions)
- [The Weekly Re-Audit](#the-weekly-re-audit)

## The Output Equation

A manager's output is the output of their organization plus the output of the neighboring organizations under their influence. Nothing on the calendar has value in itself — a meeting, a review, an approval matters only through the output it eventually creates or destroys. Leverage is the exchange rate: output created per unit of managerial time spent on an activity.

Three ways to raise output follow directly:

1. **Speed up** — do the same activities faster (helps a little, caps quickly).
2. **Raise the leverage of existing activities** — better preparation, better timing, bigger audiences for the same hour.
3. **Shift the mix** — replace low- and negative-leverage activities with high-leverage ones. This is where almost all of the gain lives, and it is what the audit below finds.

The audit takes one logged week and roughly ninety minutes of analysis. Most managers who run it discover that 30-50% of their week is spent on activities that either anyone could do, that nobody should do, or that actively subtract output.

## Step 1: Log a Real Week

Log five working days at 30-minute granularity. Rules that keep the data honest:

- **Log as you go, classify later.** Classifying while logging biases what you record.
- **Log what actually happened**, not what the calendar said. The 9:00 "deep work block" that became forty minutes of Slack is logged as Slack.
- **Mark interruptions with a tally**, not a block. Six pings inside one hour is its own finding.
- **Note who else was present** for every meeting — you will need headcount for cost math later.
- **Pick a typical week.** Not launch week, not the week after reorg. If no week is typical, log two.

A spreadsheet with four columns is enough: time, what, who, interrupt count.

## Step 2: Classify Every Block

Grove's insight is that all managerial activities reduce to a small set. Tag each block with one of:

| Type | What it looks like | Notes |
|------|--------------------|-------|
| **Information gathering** | 1:1s, reading reports and dashboards, hallway/Slack conversations, customer calls, reading code or tickets | The base activity — everything else depends on its quality. Verbal sources are fastest but vaguest; written reports discipline the writer more than they inform the reader |
| **Information giving** | Announcements, briefings, documentation, answering questions, setting context in meetings | Includes conveying not just facts but objectives, priorities, and preferred ways of doing things |
| **Decision-making** | Choosing vendors, approving designs, allocating people, setting priorities — or participating in someone else's decision | Split "made the decision" from "sat in a meeting where a decision happened to me" |
| **Nudging** | Suggesting a direction without ordering it: a comment in a design review, a question in a planning doc | Legitimate and distinct from deciding — you nudge many times a day |
| **Being a role model** | Visible behavior: how you run meetings, handle incidents, treat people, write | Values transmit through observed behavior, never through speeches. You are doing this all day whether you intend to or not |

Anything that fits none of these — doing IC work, expediting a ticket, formatting a slide — gets tagged **doing**, and becomes a delegation candidate in Step 3.

## Step 3: Score the Leverage

Now score each block high, neutral, or negative. The test for each:

- **High leverage:** one hour affects many people's output (training, a well-run staff meeting, hiring), affects one person's output for a long time (a prepared review, a career conversation, an early spec read), or affects a large body of work through perfect timing (catching a wrong design before the build starts).
- **Neutral:** necessary, output-preserving, low multiplication — expense approvals, routine syncs, most email.
- **Negative leverage:** the hour subtracted output from others. See the catalog below.

Then compute three numbers: percentage of week in high-leverage work, in "doing", and in negative leverage. Typical first-audit results for a new engineering manager: 15% high, 35% doing, 10% negative, the rest neutral. Target after redesign: 40%+ high, under 10% doing, zero tolerated negative.

## The Leverage Catalog

**Reliably high-leverage activities:**

- **Training you deliver yourself** — Grove's arithmetic: four lectures taking ~12 hours of preparation, delivered to ten people who will work ~20,000 hours in the next year. A 1% improvement buys 200 hours of output for a dozen hours of work.
- **Performance reviews prepared properly** — one written assessment steers a year of one person's output.
- **1:1s** — ninety minutes can raise the quality of a subordinate's work for two weeks or more.
- **Writing once for many readers** — a decision memo, an onboarding doc, a postmortem; the alternative is explaining it eleven times.
- **Early-stage inspection** — an hour on a one-page spec saves a month on a wrong build.
- **Hiring** — few hours have a longer half-life.
- **The timely nudge** — one question in the right design review redirects a quarter of work.

**Negative-leverage activities (each hour subtracts):**

- **Meddling** — supervising in detail a person who has high task-relevant maturity for the task. The expert slows down, stops owning outcomes, and learns to wait for instructions.
- **Waffling** — sitting on a decision others are blocked on. Ten people idling for three days is a costlier purchase than almost anything you could buy with a PO.
- **Mood contagion** — a manager's visible anxiety or gloom propagates; the team spends its energy reading you instead of building.
- **The unprepared meeting** — eight people discovering the agenda live.
- **Being the approval bottleneck** — sign-offs queueing behind your travel schedule.
- **Last-minute review** — "fixing" finished work that should have been inspected at the spec stage; you pay full rework cost and demoralize the author.

## The Delegation Protocol

Delegation is how "doing" hours convert to high-leverage hours — but delegation without monitoring is abdication, and silent re-grabbing is worse. The protocol:

1. **Pick what to delegate: the tasks you know best.** This feels backwards and is not. Monitoring is cheap when you can judge the work at a glance; delegating what you understand least means you cannot supervise it at all. Your famous-for tasks are your best handoffs.
2. **Hand off outcomes, not steps.** State the output expected, the constraints (budget, deadline, interfaces), and the quality bar. Put it in writing — two paragraphs, not a contract.
3. **Agree on the monitoring plan in the handoff conversation.** Checkpoint cadence, what gets sampled, what triggers escalation. Monitoring announced up front is quality assurance; monitoring imposed after a stumble is punishment.
4. **Monitor at the task level, not the person level.** You are sampling work products — plans, PRs, dashboards — the way a factory samples incoming material, not auditing the human.
5. **Vary depth with task-relevant maturity** (table below), and loosen visibly as results come in.
6. **Never take a task back silently.** If it is going wrong, say so, raise monitoring frequency, add training — and if you must repossess it, do it explicitly with reasons. Quiet repossession teaches the team that delegation is a trap.

## Monitoring Depth by Task-Relevant Maturity

| Subordinate's TRM for this task | What you review | Cadence | Style |
|---------------------------------|-----------------|---------|-------|
| **Low** (new to task, regardless of seniority) | The plan before work starts, then work products at each stage | 2-3 times/week, scheduled | Structured: what, when, how; short feedback loops |
| **Medium** (done it with help before) | The plan plus spot-checks of in-progress work | Weekly | Mutual reasoning: what and why; they propose, you probe |
| **High** (done it well repeatedly) | Final outputs and a few agreed indicators | At milestones / monthly | Objectives only; monitor outcomes, intervene on request or on indicator drift |

Re-rate TRM whenever the task changes. The engineer who is high-TRM on service migrations may be low-TRM on their first vendor negotiation — and your monitoring must change with them, in both directions.

## Redesigning the Calendar

Treat the calendar as a production system and apply factory rules to it:

1. **Forecast the limiting steps first.** The high-leverage fixed events — 1:1s, staff meeting, planning, reviews, training you teach — go on the calendar before anything else, like the egg timer everything else offsets from. These are commitments others schedule around; they move only for emergencies.
2. **Batch similar work.** Group PR reviews into one or two daily windows, interviews into two afternoons, email into two or three passes. Every context switch is a setup cost; batching amortizes it.
3. **Create maker blocks and defend them.** Two or three half-days of focus time, treated like meetings with yourself. Your reports' maker time deserves the same defense — audit how many of their interruptions are you.
4. **Leave slack.** Keep roughly 15% of the week unscheduled. A manager at 100% utilization is a server at 100% utilization: every arrival queues, and latency explodes for everyone downstream.
5. **Say no at the source.** Capacity planning means refusing or renegotiating work beyond capacity at the moment it is offered — "yes, after the planning cycle" or "no, and here is who can" — rather than accepting it into a queue where it will silently rot.
6. **Use the cost test on every standing meeting you own.** Attendees × hours × loaded hourly rate. An eight-person hour costs about a thousand dollars; if you would not sign a purchase order for that amount for the value produced, restructure or kill it.

A redesigned week for an EM with six reports might look like: Monday — maker block AM, staff meeting after lunch, batched reviews 4pm; Tuesday/Wednesday — 1:1s in morning blocks, office hours 2-3pm; Thursday — interviews batched PM; Friday — planning/indicator review, maker block, slack.

## Managing Interruptions

Interruptions are demand arriving at random for a resource (you) that performs best in batches. Apply production thinking rather than heroics:

- **Count them first.** The week log's tally marks show who interrupts, about what, and when. Most managers find 60%+ of interrupts come from a handful of question types.
- **Export the answers.** Recurring questions become a FAQ, a runbook, a dashboard, or a trained delegate. Answering the same question eleven times is a documentation failure, not diligence. Each written answer is a small machine that works while you sleep.
- **Batch the rest with office hours.** Two predictable daily slots convert random arrivals into a queue with a known service time. Pair with the 1:1 hold list: "great question — put it on the hold list for Tuesday" is a complete, polite sentence.
- **Publish an escalation ladder.** Define what justifies an immediate page (production down, customer-visible failure, a person in trouble) versus the next office hours versus the next 1:1. People interrupt randomly when they cannot predict what you consider urgent.
- **Do not go dark.** The goal is batching, not unavailability — an unreachable manager simply trains people to make uninformed decisions in private.

## The Weekly Re-Audit

The full audit is occasional; this five-minute version runs every Friday:

1. What fraction of my hours this week were high-leverage? (Target: rising toward 40%.)
2. What did I do that someone on the team could now do — and what training or delegation would make that true next month?
3. Where did I create negative leverage — a delayed decision, a meddled task, a mood I broadcast — and what is the repair?
4. Which interrupt should become a document, a delegate, or a hold-list item?
5. Is next week's calendar forecast — 1:1s, maker blocks, batches — already in place, or will the week happen to me again?

Write the answers down; four weeks of them is its own stagger chart of whether your management is trending toward output.
references/meetings-and-one-on-ones.md
# Meetings and One-on-Ones: The Full Playbook

## Table of Contents

- [Two Species of Meetings](#two-species-of-meetings)
- [The One-on-One](#the-one-on-one)
- [Frequency by Task-Relevant Maturity](#frequency-by-task-relevant-maturity)
- [The Subordinate's Agenda](#the-subordinates-agenda)
- [The Hold List](#the-hold-list)
- [Notes, "One More Thing", and Remote 1:1s](#notes-one-more-thing-and-remote-11s)
- [The Staff Meeting](#the-staff-meeting)
- [The Operation Review](#the-operation-review)
- [Mission-Oriented Meetings](#mission-oriented-meetings)
- [Meeting-Cost Math and When to Kill a Meeting](#meeting-cost-math-and-when-to-kill-a-meeting)

## Two Species of Meetings

Grove's claim is blunt: a meeting is the *medium* through which most managerial work — gathering and giving information, making decisions, nudging — is performed. You cannot oppose meetings any more than a carpenter can oppose hammers; you can only run good ones or bad ones.

- **Process-oriented meetings** run on a regular cadence with a known format: one-on-ones, staff meetings, operation reviews. Regularity is what makes them cheap — agendas are standing, expectations shared, setup cost near zero — and what makes them effective: problems get caught while they are small.
- **Mission-oriented meetings** are ad hoc and exist to produce one specific decision. They are expensive by nature (interrupt-driven, unprepared attendees, no standing format).

The design goal: process meetings carry the bulk of managerial work — on the order of a quarter of a manager's calendar — so that mission-oriented meetings become rare. Grove's malorganization test: if ad hoc mission-oriented meetings consume more than about 25% of a manager's time, the regular process is failing. And the corollary diagnostic: **any "ad hoc" meeting that keeps recurring is a process meeting in denial** — give it a cadence, a format, and a standing agenda, and its cost collapses.

## The One-on-One

The 1:1 is a regular, scheduled meeting between a supervisor and one subordinate, and it is **the subordinate's meeting**: their agenda, their airtime, their problems. Its purposes: mutual teaching (the supervisor shares context and skills; the subordinate teaches the supervisor what is really going on below), exchange of information too awkward or early for any other forum, and the supervisor's single best source of organizational knowledge.

Grove's leverage math: ninety minutes of your time can enhance the quality of your subordinate's work for two weeks — eighty-plus working hours. There is no cheaper multiplier on the calendar, which is why "too busy for 1:1s" inverts the economics: the busier the period, the more the 1:1 pays.

Baseline mechanics:

- **Length:** an hour as the default. Anything shorter, Grove observed, makes the subordinate confine themselves to simple things; the hard topics need room to surface. 45 minutes is the floor for a working cadence; never 15.
- **Setting:** the subordinate's turf or neutral ground, not a summons to your office. Remote: cameras on, both in a shared doc.
- **Scheduling:** rolling and protected. Reschedule within the week when needed; cancellation is reserved for true emergencies, because each silent drop teaches the report their problems can wait.

## Frequency by Task-Relevant Maturity

Set frequency per person by their task-relevant maturity on their *current* work — not by seniority, tenure, or how much you enjoy the conversation:

| Situation | Frequency | Length |
|-----------|-----------|--------|
| New to role, project, or domain (any seniority) | Weekly | 60 min |
| Ramping: competent with support, task still changing | Every two weeks | 45-60 min |
| Veteran on stable scope, strong indicators | Every three to four weeks | 60 min |
| Anyone during a crisis, reorg, or major task change | Snap back to weekly | 60 min |

Re-rate at every task change. The principal engineer who just became a manager is low-TRM again; the weekly cadence returns even though they are your most senior person. When in doubt, err frequent — over-meeting costs an hour; under-meeting costs a surprise in month three.

## The Subordinate's Agenda

The subordinate prepares and sends the agenda the day before. Preparing it is itself the point — it forces them to survey their own work, spot trouble early, and decide what matters. A copy-paste template:

```
1:1 — <name> / <manager>, <date>

1. Indicators since last time: what moved, what I think it means
2. Wins and what's finished
3. Problems and blockers — current, and ones I can smell coming
   ("anything that bothers me", half-formed worries explicitly welcome)
4. Plans: what I intend to do next, decisions I'm about to make
5. Hold-list items (accumulated small topics)
6. Career / growth / feedback item (at least monthly)
7. Manager's items (last, and they don't get to eat the hour)
```

Two supervisor rules. First, items 3 and 4 are where you earn your keep: the subordinate talks, you ask the *second* question — "what's underneath that?", "what would have to be true?", "what happens if it slips?" — until the real shape of the problem is on the table. Second, your own topics ride in seat 7. The fastest way to kill a 1:1 program is to convert it into a status extraction or a tasking session; reports stop bringing problems and start bringing performances.

## The Hold List

Both parties keep a running list — paper, doc, or DM-to-self — of topics that are important but not urgent, batched for the next 1:1. Effects: interruptions drop on both sides (most "got a sec?" pings are hold-list items in disguise), small items actually get handled instead of evaporating, and the 1:1 always has a floor of concrete material. "Put it on the hold list" is the polite, complete answer to most non-urgent pings — and modeling it teaches the team to batch their interruptions of each other, too.

## Notes, "One More Thing", and Remote 1:1s

**The supervisor takes the notes**, on the subordinate's agenda, during the meeting. Writing while they talk signals that what they said registered and will be acted on; the written trail forces follow-up on both sides. End by reading back the actions: who does what by when. Next meeting opens with that list.

**"One more thing."** When the agenda is exhausted and chairs start to push back, ask: *"What else is on your mind?"* — and wait through the pause. The genuinely important item — the resignation being considered, the conflict with a peer, the doubt about the project — surfaces disproportionately in the last five minutes, after trust has warmed and the formal agenda no longer protects it. Budget for this; ending a 1:1 ten minutes "early" without asking is leaving the most valuable item in the room.

**Remote 1:1s.** The agenda doc is open on both screens and notes are typed into it live — Grove's "both looking at the same paper", updated. Cameras on if at all possible; notifications off (a glance at Slack reads as "you don't matter" at distance). Because hallway leakage is zero remotely, the hold list and the "one more thing" question carry even more of the load — skip neither.

## The Staff Meeting

The staff meeting is the supervisor plus all direct reports, weekly, 60-90 minutes. Its special value: it is where the manager watches and shapes how their reports interact with *each other*, and where issues touching two or more people get settled with everyone present.

The format Grove prescribed is **controlled free discussion**: more structured than a bull session, looser than a briefing. The manager's role is moderator — Socratic prodder — not lecturer. If the manager talks more than a quarter of the time, it has become a broadcast, which email does cheaper.

Copy-paste agenda:

```
Staff meeting — <team>, weekly, 60 min

1. (5)  Metrics minute: the team's 4-6 indicators on screen; anomalies flagged only
2. (10) Round: one minute each — top priority, top worry. No narration of the week
3. (35) Two or three debated topics, flagged in advance by anyone on the team:
        decisions touching 2+ people, cross-cutting risks, contested priorities.
        Owner states the question; manager moderates; end each with
        decision / owner / date, or an explicit "decide by <date> with <input>"
4. (5)  Announcements that genuinely can't be written down
5. (5)  Actions read back
```

Rules: anything that concerns only one person moves to that person's 1:1; status that can be written, is written; "open issues" without an owner do not survive the meeting. The debated-topics slot is the heart — it is where the manager teaches decision-making by moderating instead of pronouncing.

## The Operation Review

The operation review is the periodic formal forum — monthly for a department, quarterly for an org — where line managers present their indicators and trends to people who do not see their work daily: adjacent teams, skip-level management, newer employees. It teaches, it cross-pollinates, and it motivates: work that will have an audience gets finished differently.

Four roles, each with a job:

- **Organizing manager:** owns logistics, picks presenters, enforces the clock, circulates pre-reads 48h ahead.
- **Reviewing manager** (the senior person): sets the tone, asks the questions juniors won't, sanity-checks trends — and *never ambushes*. Their most important behavior: visibly rewarding honest bad news.
- **Presenters** (2-4 line managers/leads): show trends and forecasts, not snapshots; show indicator pairs together; state misses before being asked; bring one problem and one ask.
- **Audience:** obligated to engage — questions are the mechanism. A silent audience converts the review into theater.

Copy-paste agenda (75 min):

```
Operation review — <org>, monthly

1. (5)  Reviewing manager: context; follow-ups from last review
2. (3 × 15) Presenters: indicator trends vs forecast (stagger charts),
        one problem being worked, one ask of the room
3. (15) Cross-area questions and discussion
4. (10) Reviewing manager: themes, decisions, actions with owners
```

## Mission-Oriented Meetings

When an ad hoc meeting truly is needed, it exists to produce a decision, and it has a **chairman** — the person who called it — with non-delegable duties:

- Name, in the invite, the decision to be made and by when.
- Invite the necessary people only — for a decision discussion, six or seven is the ceiling; beyond that, people watch instead of deciding.
- Send the pre-read 24h ahead; open by restating the question, not by reading the deck.
- Close with the decision stated in one sentence, plus owner and date — or an explicit deferral naming what information will settle it and when.
- Send the one-paragraph summary the same day.

If the meeting could not state its decision in the invite, it was an information meeting wearing a costume — handle it in writing or fold it into a process meeting. And if you are scheduling substantially the same "ad hoc" meeting for the third time, stop: give it a cadence and a standing format, and it becomes cheap.

## Meeting-Cost Math and When to Kill a Meeting

Grove's framing: a manager's time has a real dollar value, and calling a meeting is spending it — an unplanned meeting is an unplanned purchase order. Make the math explicit:

```
cost = attendees × duration (h) × loaded hourly rate
```

Ten people × one hour × $130 loaded ≈ $1,300 — roughly a new laptop per occurrence, $67k per year if weekly. The test for any standing meeting: *would you sign a purchase order for this amount, for what this meeting produced last month?*

Kill, shrink, or convert when:

| Signal | Action |
|--------|--------|
| No decision, surfaced problem, or changed plan in the last 3 occurrences | Kill it; replace with a written update |
| Half the room never speaks | Shrink the list; minutes to the rest |
| Same two people do all the talking | Make it their 1:1 or a working session |
| Purely one-way status | Convert to a dashboard or memo; reclaim the hour |
| "We've always had it" is the only defense | Cancel for a month; reinstate only what is actually missed |

Run a quarterly meeting audit: list every standing meeting you own, compute each one's annual cost, and re-justify it against what it produces. Default for borderline cases is cancellation — a genuinely needed meeting announces itself by the problems that appear in its absence, and it can always be reinstated leaner.