,

The Missing Return Path

Photo by karina trinidad on Unsplash

Why organizations fail in slow motion, and in full view of everyone who works there.

Almost everyone who has worked inside an organization shares this story: a project is doomed, a policy is going to create a crisis, an ill-advised product is going to fail spectacularly. And ‘everyone’ knows it.

The problem, when the fateful day arrives, even arrives with a response everyone saw coming; urgent memos, visits from the head office, testy phone calls. What’s strange isn’t that nobody saw it, it’s that ‘everyone’ did! But seeing it didn’t change anything.

This isn’t a knowledge problem (the information exists). It could be a vantage problem: the people near the problem are the only ones who can see it. It could be a scope problem: a strategist making decisions they’ll never see the impact of. It could even be a tempo problem, a quarterly check-in making decisions with immediate consequences they won’t see until the next quarter. Either way, the solution is a path.

The organization needs a low-impedance route, for the feedback signal to travel from where it is visible, to where it can be acted upon. An organization can get plenty of things right (clearly defined purpose, sensible boundaries, real and effective autonomy) and still experience this type of failing. None of those things permit a challenge of the system itself.


Why the signal doesn’t travel on its own

Four well-documented dynamics explain why organizations reliably fail to act on things “everyone” can see.

Authority and knowledge sit in different places. The people who can see a problem most clearly are usually not the ones with the power to fix it. Every layer, between the authority-to-act and visibility-of-the-problem, mutes the signal as it passes through the system. By the time a concern reaches someone with authority-to-act, it arrives faintly, appearing less impactful, and easy to dismiss.

Responsibility diffuses when many people know they all have the same information. Social psychologists Bibb Latané and John Darley’s classic research on bystander behavior found that individuals are less likely to act on a problem, the more other people are aware of it; each assumes someone else, probably someone better positioned, will step in. In an organization, this can show up as an entire team watching the same slow failure, each member sure that it isn’t their job to be the one who says so.

The problem stops being a problem. Sociologist Diane Vaughan’s study of the Space Shuttle Challenger disaster gave this a name: the normalization of deviance. The engineers at Morton Thiokol had documented a problem (O-ring erosion) on prior flights for years. Every time the shuttle flew safely despite it, the anomaly became a little less anomalous. The night before the Challenger launch, engineers raised the alarm, but they were overruled. The signal existed. It simply had no channel with enough authority-to-act, or with a low-enough tolerance for anomalies.


What a working return path looks like

The clearest real-world design for this problem comes from a factory floor. Toyota’s andon system gives any worker on a production line the authority to flag problems and defects as they see them, no permission required. Pulling the cord sounds an alert and starts a short response window (often measured in seconds) during which a team leader immediately responds to the workstation. If the problem is resolved before the line reaches a certain point downstream, production never even stops; if it isn’t, the line halts there until the issue is resolved.

It’s a designed, low-friction path from where a defect is first visible to where a response is guaranteed, bypassing every layer of hierarchy in between.

Four features allow this system to work. Each is a separate design decision an organization has to make.

Authority is placed at the point closest to the signal, not routed through the chain first. The person who sees the defect acts on it. A permission requirement (ask before you act) is replaced with a bounded constraint that guides the action; with no approval step in between. A permission requirement routes the decision through a person with authority; the constraint hands the decision to the person inside a boundary.

The action is cheap, specific, and protected. Pulling the cord isn’t treated as a performance event or an escalation with career risks; it’s the expected behavior, and it’s actively encouraged. Reports on Toyota plants have described pull rates in the thousands per week at some facilities. This is not the sign of a factory in crisis, but a working system demonstrating how much would have otherwise been missed.

Response is fast and cheap. The entire mechanism trusts the judgment of the worker, and that someone will reliably show up, quickly, every time. Without the guaranteed, fast response the system falls into signal fatigue.

It’s scoped narrowly. The cord exists for a specific class of event, a defect on the line. That narrowness is what keeps the signal undiluted.

W. L. Gore’s waterline principle is related. At Gore, associates are expected to consult with colleagues that are knowledgeable about the matter, before taking any action serious enough to threaten the company. “Below the waterline” decisions, to use the company’s own phrasing, are otherwise free to be acted on without asking permission first.

The clarity about when to consult (before a serious decision) means objections are raised ahead of time, never after. Where andon is a fast-response channel for something already visibly wrong, the waterline principle is a checkpoint for dissent before a risk is taken. A way to flag what’s wrong now, and a guideline before something goes wrong.


Architecting a return path

Step 1: Decide what class of signal this path is for. Don’t build one channel for everything. Andon exists for a specific kind of event, at a specific point in the process. Decide what deserves this treatment and build an appropriately scoped path.

Step 2: Give authority to the people closest to the signal. Design the action: who to contact, through what channel, with what response. Avoid the hierarchy that’s most likely to be the source of the problem. If the return path goes through a manager contributing to a problem or dynamic, the return path gets closed, or the signal gets muted.

Step 3: Make raising it cheap and protected. A policy that says retaliation won’t happen is not the same as people believing it. Amy Edmondson’s research on psychological safety (later echoed in Google’s internal study of what separated its highest- from lowest-performing teams) found that teams where people felt safe raising concerns didn’t have fewer problems; they surfaced more of them, because raising a problem wasn’t personally risky.

Step 4: Attach a guaranteed, fast response. Set a deadline and name an owner.

Step 5: Close the loop back to the person who raised it. Report back on what happened. Silence teaches the same lesson as a slow response: don’t bother next time.

Step 6: Monitor the health of the channel itself. A channel that’s never used deserves the same suspicion as a metric that’s always green. It’s just as likely to be evidence the channel isn’t trusted, as evidence that nothing is wrong. At Toyota, extremely high, sustained pull rates, were treated as a sign the system was working as intended, not as a crisis to suppress.


Problems and pitfalls to avoid

Fear doesn’t go away because a policy says it should. Even with strong protections, we go by what we see, not what we’ve been told. The first person who’s punished or ignored, quietly, or even just seems to have been, teaches everyone that the channel is a trap. No simple reassurance can repair that.

Overuse in the early period is normal. When a return path first works, expect a surge of signals as people test whether it’s real and clear a backlog of things they’d been sitting on. This is the system calibrating, not a symptom of a problem.

A channel that isn’t wired to real power doesn’t fix anything. If the person who receives the signal has no authority over what’s being flagged, the mechanism becomes theater: people take the personal risk of speaking up, nothing changes. Building the channel sometimes forces an uncomfortable question: does anyone have the power to act on this kind of problem? If the answer is no, then that has to be fixed before the channel will mean anything.

Response ownership becomes a new bottleneck. If only one person or one small team owns the response, there is a risk of creating the resistance problem this system was built to avoid. This needs dedicated capacity built around how often the channel is actually used.

Middle management will experience this as bypassing their authority.

It does, by design. Left unaddressed, this may create a legitimate source of resistance to the system itself. Enroll management in the change. Most communication should still flow through ordinary channels; this one exists for specific cases.


Further reading

  • Vaughan, Diane. The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA (1996) — the original account of normalization of deviance.
  • Latané, Bibb, and John M. Darley. The Unresponsive Bystander: Why Doesn’t He Help? (1970) — the foundational research on diffusion of responsibility.
  • Edmondson, Amy C. “Psychological Safety and Learning Behavior in Work Teams.” Administrative Science Quarterly, 44(2), 1999; and The Fearless Organization (2018).
  • Ohno, Taiichi. Toyota Production System: Beyond Large-Scale Production (1988) — the original source on jidoka and andon.
  • Liker, Jeffrey, and David Meier. The Toyota Way Fieldbook (2005) — a detailed account of how andon response actually works in practice.