Before · person-to-person
- Recipient in class — waiting
- Recipient off campus — waiting
- Recipient never opens the app — expired
Every branch parks the sender in a state they cannot resolve, during the only free minutes they have.
One autonomous system. Two radically different experiences.
Three nodes orbiting one core — the identity was built around the actual product shape: a fleet in motion around a single coordinating layer.
In short
A robot that drives itself across campus still needs two humans to succeed: someone who trusts it enough to hand over a package, and someone who can tell within seconds whether the fleet is fine. Those two people need almost nothing in common, and building one interface for both is the most common way this category fails.
The problem
Campus delivery fails on coordination, not distance. And the system that solves it produces two irreconcilable users — a student with four minutes between classes, and an operator responsible for a dozen autonomous machines at once.
What I owned
The outcome (honest)
A designed system and a validated interaction model — not an operating fleet. No robot ran on this design. What exists is a product boundary, two complete experiences, an AI authority model, and a measurement plan for the claims I could not yet test.
There are no fleet KPIs on this page, because I did not own a fleet.
The box I was designing inside
A closed campus is not a small city. The boundaries are tighter, the identities are known, and the time budget is brutal — which rules out most of what consumer delivery apps do.
Problem & stakes
A ten-minute walk across campus is not a ten-minute task. It becomes a forty-minute interruption once you add remembering, finding a window, locating a person inside a building, waiting for them to show up, and getting back late. The walk is the only part anyone would have estimated.
The distance is unchanged. What collapses is the number of moments where two people have to be synchronised — from five down to one, and that one happens on the sender’s schedule.
Strategy pivot
My first flows had the sender pick a person, then wait for that person to accept. I spent a week designing the waiting screen — pending states, nudges, expiry timers, a graceful way to fail. All of it was craft applied to a step that should not have existed. Inside a five-minute class break, a pending state is a dead end with good typography.
Before · person-to-person
Every branch parks the sender in a state they cannot resolve, during the only free minutes they have.
After · department pre-acceptance
Departments accept once, as a policy, instead of individuals accepting every time. The recipient stops being a blocking dependency.
What this decision cost
Pre-acceptance is an organisational commitment, not a feature. It needs a department to agree that a robot may leave a package at its desk during posted hours, and someone there to own what arrives. That is a policy conversation I could scope but not close — so the design has to work when a department says no.
How the design absorbs a “no”
Person-to-person sending stays as a first-class path with an honest constraint attached: it is offered when the system has signal that the recipient is reachable, and it carries a visible risk of return. The pivot changes the default, not the option set.
Accountability
This was design work wrapped around an autonomy stack I did not build and a university policy surface I did not control. Being precise about that boundary matters more than sounding senior about it.
I owned
I did not own
Decisions & tradeoffs
Each of these had a credible alternative that a reasonable designer would have picked, and each one cost something real. The costs are listed because they were accepted knowingly, not discovered afterwards.
User experience
Decision
The user-facing product exposes one machine fact — where the package is and when it arrives. Telemetry, robot identity, fleet state and every other robot on campus stay out of the interface entirely. The user tracks a delivery, not a machine.
Alternatives considered
Show the fleet map with all robots, which is what every robotics demo does because it looks impressive. Or expose battery and sensor confidence as a transparency feature, on the argument that autonomy earns trust by being legible.
Why I chose it
Transparency only builds trust when the person can act on what they see. A student cannot do anything with “localisation confidence 91%” except worry. Exposing it transfers operational anxiety to someone with no authority to resolve it — the worst trade in the system.
Cost accepted
When something goes wrong, the user has no independent way to verify what happened, so the exception copy has to do all the work — and if it ever lies, the whole abstraction collapses at once. I also gave up the demo appeal of a live fleet map, which is the single thing stakeholders ask for first.
Operations
Decision
Command opens on what needs a human, ranked, with the map as supporting context rather than the primary object. Twelve healthy robots produce one line of text; one at-risk robot produces a card with a recommendation attached.
Alternatives considered
The conventional fleet console: a full-bleed live map where every robot renders identically and the operator scans for anomalies. Or a dashboard-first layout leading with utilisation and throughput charts.
Why I chose it
An equal-density map makes the operator’s job pattern-matching against a moving picture, and that degrades exactly when the fleet is busiest. The target was five seconds to answer four questions: is anything wrong, which robot, why, and do I need to act. A ranked queue answers all four; a map answers none of them without interpretation.
Cost accepted
The queue is only as good as the ranking behind it, so a missed signal is now invisible rather than merely hard to spot — a map at least gives the operator a chance to notice something the system didn’t flag. Experienced operators also lose the ambient spatial awareness they build from watching a map, which is real skill I am deliberately trading away for triage speed.
AI behaviour
Decision
The system automates only reversible, low-consequence actions — a minor reroute around congestion, a charging decision for an idle robot. Anything that changes what a user was promised, moves a mission between robots, or touches a stopped robot requires an explicit human approval with the expected consequence stated before the click.
Alternatives considered
Full autonomy with exception-only escalation, which is where fleet products eventually want to go and where the operational savings are. Or the opposite: advisory-only AI that never acts, on the grounds that oversight is safer when nothing is automatic.
Why I chose it
Full autonomy asks for trust the system has not earned on day one, and the first bad automatic decision would end operator confidence permanently. Advisory-only is the reverse failure — it buries the operator in confirmations for things nobody would ever say no to, which is how approval becomes reflex. Splitting on consequence and reversibility keeps approval meaningful in the moments where it actually is.
Cost accepted
A human in the loop is a latency floor: incidents wait for someone to look. The model also does not scale — it holds for a campus fleet and breaks somewhere past the point where one operator can read every recommendation. Worst of all, the boundary between “reversible” and “consequential” is a judgement I encoded, and a wrong call there is invisible until it causes an incident.
Product strategy
Decision
Remove the acceptance handshake from the critical path. Departments carry standing acceptance during posted hours, and for person-to-person sends the system predicts reachability from class patterns, past acceptance and building access — then warns before confirm rather than failing afterwards.
Alternatives considered
Keep the handshake and make waiting pleasant: live status, nudges, a generous expiry. Or schedule everything in advance so both parties commit to a slot up front.
Why I chose it
Both alternatives solve the wrong problem. A better waiting screen still consumes the class break, and scheduling reintroduces the calendar negotiation the product exists to eliminate. Prediction moves the uncertainty to where it is cheap — before the user has committed anything — and leaves them with a choice rather than a pending state.
Cost accepted
The system now makes a claim about a person’s availability, which is a claim it can get wrong and which touches privacy in a way an acceptance tap does not. I kept the signals coarse and campus-scoped for that reason, and accepted weaker predictions as the price. There is also a cold-start hole: a new recipient has no history, so the first send to anyone is the least informed one.
One system, two humans
Same robot. Same minute. Same mission. Below is Mission #482 rendered for both people — the clearest test of whether the two-product boundary actually holds, because if one side needs something from the other’s screen, the split was wrong.
VYN User · iOS
One fact, one place, one action if needed. No robot identity, no fleet, no telemetry.
Live delivery
Your robot is
6 minutes away
Handoff code
4 8 2 1
VYN Command · Web
Fleet state compressed to one rail, and the single thing that needs a person next to the action that resolves it.
Needs attention 1
Recommended
Transfer #482 to VYN-03 — idle, 180 m away, 94% charge.
Consequence: ETA 6 min → 8 min. Under the 3-minute threshold, so the recipient sees no change.
11 other robots nominal · 7 missions on schedule · no operator action required
Both interfaces are rebuilt in HTML and CSS from the design files rather than pasted in as flat exports, so the type, states and data stay legible at any width. The values shown are the worked example I designed against.
AI product rigor
An operations AI is not a feature you add to a dashboard — it is a set of standing permissions. The design work was deciding where each of those permissions stops, and what the product does on the day the model is confidently wrong.
The escalation ladder
01
System
Detects the condition and states it plainly. No interpretation yet.
02
System
Adds the signals behind it, so the operator can check the reasoning rather than trust it.
03
System
Proposes one action with its expected consequence and at least one alternative.
04
Human
The gate. Anything consequential stops here until a person accepts, modifies or refuses it.
05
System
Only reversible, low-consequence policies run unattended — and every one of them is logged where the operator reads it.
06
Human
Safety, property or a stopped robot. The system stops proposing and simply hands over.
What an operator sees before approving anything consequential
Authority model
The last row is the one I argued hardest for. The user-facing abstraction from Decision 01 only survives if no machine can quietly change what a person was promised.
Failure model
Impact
No robot ran on this design, no fleet was operated, and no delivery was completed. I would rather be precise than impressive, so this section separates what the work produced, what a small amount of testing supported, and what I would have to measure before defending any of it.
01 — Design outcomes
The concrete deliverables: a defended two-product boundary with a shared mission object, two complete experiences including failure paths, an attention-queue pattern for fleet triage, an AI authority model that maps actions to who decides, and a tablet deployment flow that lets a non-specialist bring a robot online without touching motor-level configuration.
02 — What validation showed
I ran informal walkthroughs with campus students and department front-desk staff — not a study, and not enough people to claim a result. What I was checking was comprehension: could a sender predict what happens next, and could a desk staffer say what a robot arriving at their counter would require of them.
03 — What I would measure first
Each of these exists to catch a specific way one of my decisions could be wrong. If I were handed a pilot tomorrow, this is the instrumentation I would ask for before anything else.
How to read this page
Every figure above describes the design or a walkthrough observation. None of them is an operational metric, a pilot result, or a production outcome — because I did not own instrumentation, and inventing the numbers would make the honest parts of this case study worthless too.
Reflection
I started on “how should a robot delivery app work?” and finished on “how should humans interact with an autonomous system when they carry completely different responsibilities for it?” Everything good in this project came after that second question replaced the first — including the decision to stop designing one product.
The four rules the system ended up running on