The escalation that gets picked up instead of handed back
Can I hand this to an engineer without them needing to gather the evidence again?
The ticket
INTERNAL: your escalation came back, please rewrite it
From: Priya, platform team lead Re: Halden Freight, this morning Impact: the fix is in, the cause is not addressed
I am sending this back rather than triaging it, and I want to explain why so it does not happen again.
I cannot tell from your note what you actually observed. You have told me there was a configuration problem, which is a conclusion, not evidence. I do not know what you saw, what you checked, or what you already eliminated, so the first thing anyone here would do is repeat all of it.
I also cannot tell what you want. "Take a look at how this got out" is not something I can assign to anybody. Tell me what you are asking for and I will tell you whether it is realistic.
Everything you need is in the notes from the incident. Please send it again.
Your job
- Rewrite the escalation so an engineer can pick it up without gathering the evidence a second time.
- Quote what you observed rather than what you concluded.
- End with a request specific enough to be assigned to somebody.
Working notes
There is no system to bring up. tse start puts the incident notes and your
returned draft in front of you.
Rewrite the draft, then run tse check. This is graded on nearly the opposite
of the customer update you wrote last time, and that is deliberate: softened
language is a defect here, and the specifics are the whole point.
- Track
- Customer communication
- Time
- about 30 minutes
- Difficulty
- Involved
- Tier
- Core
Do these first: The update you owe a customer once the outage is over, Every health check is green and the customer still gets nothing
Start it
In a Codespace or a local clone:
tse start communication/02-the-escalation-that-does-not-come-backThat provisions the broken system and prints the ticket above. Investigate with ordinary tools, then run tse check.
Investigation scratchpad
Saved in this browser as you type. Nothing is uploaded. 0 of 7 filled in.
In their words, not yours. Include scope and urgency.
Before running anything: target layer, expected output, two likely causes.
The command or query, and why it is safe to run here.
Three separate lists. This is the step people skip.
One proof sentence, one safe next step, one alternate hypothesis.
Plain language. Impact first. No blame, no speculation.
One gap, one command to repeat tomorrow, one confidence score.
Hints
Each hint gives away a little more. Try to spend a few minutes on your own evidence first, because the recall is what makes it stick.
Hint 1 of 3
Priya told you exactly what was wrong with it, which is more than most people get. Read her note again as a specification rather than as a complaint.
She said two things. She cannot tell what you observed, and she cannot tell what you want.
The first is the harder one to fix, because the draft does not look like it is missing anything. "It turned out to be a configuration problem" feels like an answer. It is a conclusion, and a conclusion is the one thing she cannot use: she has to either take it on faith or go and re-establish it, and she is not going to take it on faith.
The distinction worth holding on to:
What you saw is evidence. What it meant is a conclusion. An escalation needs both, clearly separated, in that order.
The second is easier and people still skip it. Something an engineer can be assigned has a verb and an object. "Look into it" has neither.
One more thing. You are writing to a colleague, not a customer. Everything you carefully removed from the update last time belongs in this one.
Hint 2 of 3
Six labels, because each one is a question the receiving engineer has to have answered before they can pick this up:
| Label | The question it answers |
|---|---|
| Impact | Who is affected, how badly, since when, and who found it |
| Evidence | What you observed, quoted rather than summarized |
| Confirmed | What you established as working, so nobody re-checks it |
| Ruled out | What you eliminated, so nobody re-checks that either |
| Suspected cause | What you think it is, marked clearly as a belief |
| Request | The specific thing you want somebody to do |
evidence.md has the observations under ## Facts, in the form they were
actually seen. Quote at least two of them verbatim. Paraphrasing an observation
turns it back into a conclusion, and then Priya is where she started.
The one people get wrong is Request. There are two separate asks available in this incident and only one of them is obvious:
- The fix is already in, but the same edit may exist elsewhere. One misconfigured service is an incident; a shared template is the next outage.
- Nothing you own could see this. Every check ran inside the thing it was checking and passed for three hours while the service was unreachable.
The second is the one worth an engineer's week. The first is worth their afternoon. Ask for both, in that order of importance, and say why.
Hint 3 of 3
Run the grader on the returned draft and read what it names:
tse check
It fails on all five. Work down them. The skeleton, with the two rules that catch most rewrites marked:
Impact: who, how badly, since when, and that a customer found it
Evidence: quote at least two observations verbatim <- the specifics rule
Confirmed: what you established as working
Ruled out: what you eliminated <- has its own rule
Suspected cause: what you think it is, marked as a belief
Request: the specific thing you want done <- ends the message
For the evidence line, the two that carry the whole diagnosis are the published mapping and the port the application announced on startup. Neither means anything alone, and side by side they are the entire finding. Quote both:
127.0.0.1:8100->8081/tcp
server_started port=8080
For the ending, the grader is looking for a commitment or a request in the last few lines. "Please confirm whether the same edit exists elsewhere" satisfies it. So does asking Priya to tell you what is realistic, which is also the honest thing to ask given you have already given the customer a date.
Around 250 words. Longer than the customer update, and it should be: nothing here is being spared the detail.
Solution
Write your customer update before you read this. Comparing your wording against the model answer is worth more than reading it cold.
Reveal the solution
Solution: the escalation that does not come back
What the evidence proved
Nothing to investigate again. The incident was mixed/01 and the notes were
handed to you. What was under test is whether an engineer can act on what you
wrote.
| Rule | What it is really asking |
|---|---|
| Answers every question an engineer will ask | Can this be picked up, or does it need a reply first |
| Quotes the technical specifics | Did you send observations, or conclusions |
| Names something ruled out | Will somebody waste an hour re-checking what you already eliminated |
| Ends with a request and an owner | Is there anything here to assign |
| Is the right length | Is the detail actually present |
Note what is missing compared to communication/01. There is no rule against
internal vocabulary here, and the rule about specifics is its mirror image. The
same module grades both, driven by the exercise's own evidence file, which is
the only reason one linter can hold two nearly opposite standards.
Root cause
The returned draft was a customer update sent to an engineer.
Every instinct that made the last exercise's message good made this one useless: it softened the specifics, replaced observations with a conclusion, and closed on a polite non-request. Priya could not have assigned it to anybody, so she sent it back, which is the cheapest possible outcome. The expensive version is the one that gets triaged, half understood, and quietly deprioritized.
Scoped fix
Rewrite labs/communication/_stack/escalation.md under six labels: impact,
evidence, confirmed, ruled out, suspected cause, request. Quote at least two
observations verbatim from evidence.md rather than describing them.
Then:
tse check
Customer update
None owed. The customer was updated in communication/01, and this escalation
is what that update committed you to. The only thing that reaches them from
here is whether you can hold the date you gave, which is why the request asks
Priya what is realistic rather than telling her what you need.
Engineering escalation, if you needed one
This is the artifact, so here it is in full:
Impact: Halden Freight, enterprise, total loss of service for all users from 06:00 to 09:12. Every request was accepted and closed with no response. Found by the customer, not by us.
Evidence: the published mapping was
127.0.0.1:8100->8081/tcpwhile the application loggedserver_started port=8080on the way up. Requests to the published address returned nothing. The container reportedUp (healthy)throughout, because the health check runs inside the container and connects straight to the application, never crossing the mapping that was wrong.Confirmed: application process health, clean application startup, database availability, and that the fault was entirely in the published mapping.
Ruled out: application crash, dependency failure, credential or data problem.
Suspected cause: last night's release changed the container-side target of the published port from 8080 to 8081. The application has always listened on 8080.
Request: two things, and the second matters more than the first.
The one-digit fix is already in. Please confirm whether the same edit exists in any other service that shares this release template, because one misconfigured service is an incident and a shared template is an outage waiting for the next deploy.
We have no signal that can see this class of fault. Every check we own runs inside the container and passed for three hours while the service was completely unreachable. A check that connects from outside the published address would have caught this in seconds. Can the platform team own adding one, and tell me what is realistic, so I can hold to the date I gave the customer.
The second request is the one that makes this escalation worth writing. The first closes an incident. The second closes a class of incident, and it is available to you only because you noticed that every signal was green.
Check your understanding
Three questions on what the evidence here proved, and what it pointedly did not. Wrong answers explain themselves, and so do right ones.
tse quiz
Check your understanding
Three questions on what the evidence proved and what it did not. Every answer explains itself, including the right one.
3 questions, none answered yet.
Why this one exists
An escalation is graded on nearly the opposite of a customer update. Softened language is a defect here, and the technical specifics are the point. What makes one get picked up is a request that names something an engineer can actually do.
In an interview
The fastest way to lose credibility with engineering is a stream of escalations that have to be sent back for detail. The fastest way to gain it is one that arrives with the evidence attached and a specific request at the end. Interviewers ask for this in writing because it is the artifact that shows whether somebody has actually worked a queue.
Commands introduced
tse check
Evidence layers
- the written escalation