Virtual Dojo

Virtual Dojo lets a coach review fight footage with athletes who aren't in the room. Rather than routing video peer-to-peer, it runs a mediasoup Selective Forwarding Unit so a session scales past two participants without every client uploading a separate stream to everyone else. The footage itself is served from object storage and played locally by each participant, while the server relays only playback intent — play, pause, seek — keeping everyone on the same frame without re-encoding video. Annotation strokes are broadcast in normalised coordinates so a mark drawn on a phone lands in the same place on a desktop.

Tech Stack

Next.jsReactNode.jsmediasoupWebRTCSocket.IOMongoDBCloudflare R2AWS EC2Tailwind CSS

Key Features

  • mediasoup SFU for multi-party audio/video instead of peer-to-peer mesh
  • Coach-authoritative playback sync — play, pause and seek mirrored to every athlete
  • Real-time drawing annotation over the footage, shared live between participants
  • Resolution-independent strokes via normalised coordinates
  • Stable identity across reconnects, so a dropped athlete rejoins in the same role
  • Single-session enforcement — a second tab evicts the first cleanly
  • JWT auth over httpOnly cookies with rate limiting
  • Presigned uploads to Cloudflare R2 for footage storage

Why Virtual Dojo uses a mediasoup SFU

Each participant sends one camera stream. Change the inputs to compare upload demand.

Participants6
Camera stream1.5 Mbps
Peer-to-peer mesh
Every participant uploads to every other participant.
Peer-to-peer mesh topologyP1P2P3P4P5P6
Upload per participant
7.5 Mbps
Total upstream
45.0 Mbps
Connections per client
5
mediasoup SFU
Each participant uploads once; the server forwards streams.
mediasoup SFU topologySFUP1P2P3P4P5P6
Upload per participant
1.5 Mbps
Total upstream
9.0 Mbps
Connections per client
1

System Design

SFU Instead of a Peer-to-Peer Mesh

In a mesh call every participant uploads a separate stream to every other participant, so upstream bandwidth grows with each person who joins — fine for two people, unusable for a coach and a squad. Virtual Dojo runs a mediasoup Selective Forwarding Unit on the server instead: each client uploads its stream once, and the SFU forwards it to everyone else. Adding a participant costs the client nothing extra, and the server forwards packets without decoding or re-encoding them.

Playback Sync Without Restreaming the Footage

Fight footage is stored in Cloudflare R2 and played locally by each participant, so the video never travels through the SFU. The server only relays playback intent — play, pause and seek events over Socket.IO — and every client applies them to its own player. The coach is authoritative, so everyone lands on the same frame while the video itself stays a direct download from object storage.

Resolution-Independent Annotation

Drawing strokes are broadcast as normalised coordinates rather than pixels, so a mark is stored relative to the video frame instead of the device that drew it. A circle drawn on a phone lands in exactly the same spot on a coach’s desktop, at any window size, without rescaling logic on the receiving end.

Session Identity and Reconnects

Participants keep a stable identity across reconnects, so an athlete who drops out rejoins in the same role rather than appearing as a new person mid-session. Single-session enforcement means opening a second tab cleanly evicts the first instead of leaving a duplicate ghost participant in the room.

Auth and Media Uploads

Authentication uses JWTs delivered over httpOnly cookies with rate limiting on the auth endpoints. Footage is uploaded straight to Cloudflare R2 through presigned URLs, so large video files never pass through the application server — the API only issues the credential and records the resulting object.

Skills Demonstrated

mediasoup SFUWebRTCSocket.IOReal-time State SyncCloudflare R2JWT AuthAWS EC2

What Was Hard

The decisions that took the longest to get right.

Keeping Everyone on the Same Frame

The problem

A coach saying “watch the left hand here” is worthless if the athlete is two seconds behind. The obvious fix — stream the coach’s screen — means re-encoding video and paying for it in quality and latency. Letting each client negotiate timestamps with the others is worse: with no single source of truth, clocks drift apart and every seek turns into an argument nobody wins.

The approach

The footage is downloaded and played locally by every participant, so nobody is watching a re-encoded stream. The coach is the only authority on playback position — play, pause and seek are broadcast as intent over Socket.IO and applied by every client to its own player. There is no negotiation, so there is nothing to drift: the coach moves, everyone follows.

Annotations That Land in the Same Place

The problem

A stroke drawn on a phone is a set of pixel coordinates in that phone’s viewport. Send those pixels to a coach on a desktop and the mark lands somewhere else entirely — wrong spot, wrong scale, pointing at the wrong thing. Correcting for it on the receiving end means every client needs to know every other client’s viewport size and keep up as windows resize.

The approach

Strokes are broadcast in normalised coordinates — stored relative to the video frame rather than the screen that drew them. Each client scales them to its own player on render, so a circle drawn on a phone lands in exactly the same spot on a desktop at any window size, with no viewport bookkeeping between clients.

Surviving a Mid-Session Reconnect

The problem

Athletes train on phones and phones drop off Wi-Fi. If identity is tied to the socket connection, a five-second dropout makes the same person reappear as a stranger: a new participant in the room, stripped of their role, while their old ghost lingers in the participant list.

The approach

Identity is held separately from the transport, so a reconnecting participant is recognised as the same person and resumes in the same role rather than joining fresh. The stale connection is cleaned up rather than left behind, so a dropout costs a few seconds of video instead of a broken session.

One Person, One Session

The problem

Opening the app in a second tab produces two live participants for one human — two camera feeds, two sets of annotations, and a coach who now has to guess which one is real. Silently ignoring the second tab is no better, because the user is left staring at a page that looks broken.

The approach

Sessions are enforced as one per account: a new session cleanly evicts the previous one instead of running alongside it. The tab that loses is told why it was disconnected, so the outcome reads as a deliberate rule rather than a bug.