Built a prototype where two users can share screens simultaneously while a group watches.
WebRTC with an SFU keeps it efficient; WASM handles local transforms.


When I first sketched this out, the vision was deceptively simple: two colleagues both sharing screens simultaneously, while a third viewer watches—and everything staying smooth, low-latency and browser-native. No plug-ins. No heavy server-rendering. The combo of WebRTC + WebAssembly made that possible—and it turned into one of those “aha” projects that changed how I architect multi-party ones.

Why WebRTC + WASM?

Screen-sharing is well supported in WebRTC: you capture a display stream via getDisplayMedia and send it over RTCPeerConnection. deepstream.io+2videosdk.live+2 But the challenge was two-fold:

  • Two simultaneous shares: Each of the two presenters is sending their screen, while the viewers (and each other) may consume and even annotate.

  • Local transforms: Overlaying annotations, combining tracks (say screen + webcam), processing frames (for example cropping, highlighting) before pushing them out. That’s where WASM comes in: rather than blowing CPU in JS or relying on server-side compositing, you can do much of the heavy lifting client-side.

In short, WebRTC handles real-time peer connections and media, WASM handles local preprocessing/transform logic, and a light-weight SFU (Selective Forwarding Unit) or mesh handles multi-viewer delivery.

How I built the prototype

Here’s an overview plus a snippet of how the flow worked:

Architecture:

  • Browser clients register in a “room”. Two of them become screen sharers, others become viewers.

  • A signalling server (WebSocket/SignalR) handles offers/answers and ICE candidate exchange.

  • A lightweight SFU (hosted on my home server) forwards media streams: ensures viewers don’t get overwhelmed by peer-to-peer mesh when number grows.

  • In each sharer client: capture screen + optionally webcam, pipe through WASM module for any transformations (e.g., highlight, blur, annotate) before feeding into WebRTC.

  • Viewers receive combined feeds via SFU and can optionally send back annotations or “which screen do I want to focus on” commands over DataChannel.

Code snippet (simplified):

// 1. Capture screen
const displayStream = await navigator.mediaDevices.getDisplayMedia({ video: true, audio: false });

// 2. Instantiate WASM processing module (e.g., using Rust-WASM)
const proc = await MyWasmProcessor.init();
const processedTrack = proc.processVideoTrack(displayStream.getVideoTracks()[0]);

// 3. Create peer connection and add track
const pc = new RTCPeerConnection({ iceServers: […] });
pc.addTrack(processedTrack, displayStream);

// 4. Signalling and join room logic (omitted)

// 5. If viewer: subscribe to SFU feed, display each remote video
pc.ontrack = (ev) => {
const [stream] = ev.streams;
const video = document.createElement(‘video’);
video.srcObject = stream;
video.autoplay = true;
container.append(video);
};

Real-world relevance & context

  • WebRTC screen sharing is mature and widely supported. A tutorial from 2025 shows full plugin-free screen share workflows using getDisplayMedia. videosdk.live+1

  • But many apps still rely on server-heavy rendering or video-services. Doing more work client-side via WASM reduces server load, latency, and cost—especially useful when you want features like multi-screen share, real-time annotation, or field/craft workflows (e.g., repair centre bays) where connectivity may fluctuate.

  • Rust-to-WASM libraries exist for WebRTC/data-channels (e.g., wasm‑peers) which I experimented with for the data-channel/annotation side. GitHub

What this meant for my projects

  • I took this prototype mindset into apps like the P2P/repair workflow systems: if two users can share screen + annotate, then one technician and one assessor can collaborate live on the shop floor, even if network is spotty.

  • The “WASM handles local transforms” principle became a recurring pattern: annotate, blur sensitive parts, overlay part counts, highlight damage zones—all before upload/sync, so remote reviewers see cleaned, enriched feeds.

  • Because of this, UI latency dropped, server load fell, and the flow felt fluent to users (which in workshop/field context is big).

Looking ahead

  • In the next iteration I plan to support multi-presenter >2, automatic layout switching (who’s speaking/annotating now), and dynamic quality adjustment (WASM-side) when network degrades.

  • Add live collaborative annotations: two presenters draw/highlight on screen in real time, with local WASM sync for low latency and data channel fallback when media degrades.

  • Move more transforms into WASM + WebGL/Canvas: e.g., damage detection overlay, highlight lanes on the screen share feed, edge-detection of parts.

  • Build fallback mode: if direct WebRTC fails (poor NAT traversal), gracefully degrade to SFU with low-res video + annotation track only, keeping usability intact.

Final thought

Screen sharing is no longer “someone shows their screen, others watch”. The frontier is simultaneous shares, in-browser transforms, live collaboration, and offline/resilient modes. When you architect using WebRTC + WASM + SFU, you make that real. That’s the kind of tool-level thinking I enjoy — build forward, lean, and meaningful.