Virtual Dojo lets a coach review fight footage with athletes who aren't in the room. Rather than routing video peer-to-peer, it runs a mediasoup Selective Forwarding Unit so a session scales past two participants without every client uploading a separate stream to everyone else. The footage itself is served from object storage and played locally by each participant, while the server relays only playback intent — play, pause, seek — keeping everyone on the same frame without re-encoding video. Annotation strokes are broadcast in normalised coordinates so a mark drawn on a phone lands in the same place on a desktop.
Tech Stack
Key Features
- mediasoup SFU for multi-party audio/video instead of peer-to-peer mesh
- Coach-authoritative playback sync — play, pause and seek mirrored to every athlete
- Real-time drawing annotation over the footage, shared live between participants
- Resolution-independent strokes via normalised coordinates
- Stable identity across reconnects, so a dropped athlete rejoins in the same role
- Single-session enforcement — a second tab evicts the first cleanly
- JWT auth over httpOnly cookies with rate limiting
- Presigned uploads to Cloudflare R2 for footage storage
Why Virtual Dojo uses a mediasoup SFU
Each participant sends one camera stream. Change the inputs to compare upload demand.
- Upload per participant
- 7.5 Mbps
- Total upstream
- 45.0 Mbps
- Connections per client
- 5
- Upload per participant
- 1.5 Mbps
- Total upstream
- 9.0 Mbps
- Connections per client
- 1
System Design
In a mesh call every participant uploads a separate stream to every other participant, so upstream bandwidth grows with each person who joins — fine for two people, unusable for a coach and a squad. Virtual Dojo runs a mediasoup Selective Forwarding Unit on the server instead: each client uploads its stream once, and the SFU forwards it to everyone else. Adding a participant costs the client nothing extra, and the server forwards packets without decoding or re-encoding them.
Fight footage is stored in Cloudflare R2 and played locally by each participant, so the video never travels through the SFU. The server only relays playback intent — play, pause and seek events over Socket.IO — and every client applies them to its own player. The coach is authoritative, so everyone lands on the same frame while the video itself stays a direct download from object storage.
Drawing strokes are broadcast as normalised coordinates rather than pixels, so a mark is stored relative to the video frame instead of the device that drew it. A circle drawn on a phone lands in exactly the same spot on a coach’s desktop, at any window size, without rescaling logic on the receiving end.
Participants keep a stable identity across reconnects, so an athlete who drops out rejoins in the same role rather than appearing as a new person mid-session. Single-session enforcement means opening a second tab cleanly evicts the first instead of leaving a duplicate ghost participant in the room.
Authentication uses JWTs delivered over httpOnly cookies with rate limiting on the auth endpoints. Footage is uploaded straight to Cloudflare R2 through presigned URLs, so large video files never pass through the application server — the API only issues the credential and records the resulting object.
Skills Demonstrated
What Was Hard
The decisions that took the longest to get right.
The problem
A coach saying “watch the left hand here” is worthless if the athlete is two seconds behind. The obvious fix — stream the coach’s screen — means re-encoding video and paying for it in quality and latency. Letting each client negotiate timestamps with the others is worse: with no single source of truth, clocks drift apart and every seek turns into an argument nobody wins.
The approach
The footage is downloaded and played locally by every participant, so nobody is watching a re-encoded stream. The coach is the only authority on playback position — play, pause and seek are broadcast as intent over Socket.IO and applied by every client to its own player. There is no negotiation, so there is nothing to drift: the coach moves, everyone follows.
The problem
A stroke drawn on a phone is a set of pixel coordinates in that phone’s viewport. Send those pixels to a coach on a desktop and the mark lands somewhere else entirely — wrong spot, wrong scale, pointing at the wrong thing. Correcting for it on the receiving end means every client needs to know every other client’s viewport size and keep up as windows resize.
The approach
Strokes are broadcast in normalised coordinates — stored relative to the video frame rather than the screen that drew them. Each client scales them to its own player on render, so a circle drawn on a phone lands in exactly the same spot on a desktop at any window size, with no viewport bookkeeping between clients.
The problem
Athletes train on phones and phones drop off Wi-Fi. If identity is tied to the socket connection, a five-second dropout makes the same person reappear as a stranger: a new participant in the room, stripped of their role, while their old ghost lingers in the participant list.
The approach
Identity is held separately from the transport, so a reconnecting participant is recognised as the same person and resumes in the same role rather than joining fresh. The stale connection is cleaned up rather than left behind, so a dropout costs a few seconds of video instead of a broken session.
The problem
Opening the app in a second tab produces two live participants for one human — two camera feeds, two sets of annotations, and a coach who now has to guess which one is real. Silently ignoring the second tab is no better, because the user is left staring at a page that looks broken.
The approach
Sessions are enforced as one per account: a new session cleanly evicts the previous one instead of running alongside it. The tab that loses is told why it was disconnected, so the outcome reads as a deliberate rule rather than a bug.