Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Yes. A web page can segment a person from a webcam feed or a photo with MediaPipe’s Image Segmenter and use the resulting mask for a background blur or replacement. The phrase “entirely in the browser” needs narrower wording, though. Input frames are processed on the device, but the page still downloads model files, the API can contact Google, and GPU output has failed in specific browser versions in ways that can look correct on screen. This guide covers how to pick a model, turn its output into a mask you can render, and test failures by browser and delegate.

Choose the model by the mask you need

Start from the output rather than from a benchmark. Google AI Edge’s image segmentation guide, last updated 1 October 2026 (UTC), describes several person-related models, and each produces a different kind of mask.

Person/background selfie model

This is the right choice for a portrait background blur or replacement, because it separates the person from everything else. It is available in two input shapes: square 256×256 and landscape 144×256. The guide says landscape may be more efficient when incoming frames are consistently landscape, such as video calls. Make your crop and orientation logic match the shape you choose.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Hair-only model

Use this for hair-specific effects such as recolouring. It returns hair pixels only, so it cannot by itself show where the person ends and the background begins.

Multi-class selfie model

This model labels background, hair, body skin, face skin, clothes, and other accessories. Choose it when an effect should apply to one region and leave the others alone, for example changing clothing colour while keeping skin unchanged. On the CPU path it is the slowest model in the latency table below, so take the extra labels only when your effect needs them.

Published latency and why the delegate choice is per model

The guide reports average latency for the complete pipeline on a Pixel 6, for both the CPU and GPU delegates. These figures come from one phone model. They are not a guarantee for your users’ hardware, and they do not include your own capture, mask conversion, or compositing.

Model (as named in the guide) Mask output Pixel 6 average, CPU Pixel 6 average, GPU
SelfieSegmenter, square 256×256 Person/background 33.46 ms 35.15 ms
SelfieSegmenter, landscape 144×256 Person/background 34.19 ms 33.55 ms
HairSegmenter Hair only 57.90 ms 52.14 ms
SelfieMulticlass 256×256 Background, hair, body skin, face skin, clothes, other accessories 217.76 ms 71.24 ms
DeepLab-V3 Not stated in the guide 123.93 ms 103.30 ms

The table shows why you cannot choose a delegate once and forget it. For the square selfie model the CPU path is slightly faster; for the landscape model the GPU path is slightly faster; for the multi-class model the GPU path is roughly three times faster than CPU. The guide does not claim that one delegate wins across every model or device, so measure each model you ship on the browsers and devices you support.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Older option: the Body Segmentation API in TensorFlow.js

Many older tutorials point to the TensorFlow Blog post “Body Segmentation with MediaPipe and TensorFlow.js,” published 25 January 2022, which is linked from the TensorFlow Blog. The post documents MediaPipe and TensorFlow.js runtimes, two model types (general and landscape), and segmentation from either a video element or a still image. According to the post, general increases accuracy while reducing inference speed compared with landscape. Treat that as dated guidance: check the current package and API status before adopting it as a new dependency. The post also warns that converting a mask between representations can be expensive, so keep the format you receive where you can.

Run modes and the mask you receive

The Image Segmenter runs in three modes: IMAGE for single images, VIDEO for decoded video, and LIVE_STREAM for camera streams. A live webcam effect uses LIVE_STREAM. In that mode results do not come back from the call itself; they arrive through a result listener callback, so your render loop has to handle them asynchronously.

The segmenter can return two kinds of mask:

  • Category mask: a uint8 image in which each pixel holds a class index. It produces hard edges and is the simplest way to cut out a person.
  • Confidence masks: float values for each class. They allow soft edges that you can threshold or feather yourself.

Request only the mask type you will render. Then wire up a live pipeline in this order:

  1. Create the segmenter with the model file you chose and set the running mode to LIVE_STREAM.
  2. Enable the output you need: a category mask for hard cutouts, or confidence masks for soft edges.
  3. Register the result listener before you send the first frame.
  4. Send each camera frame with its timestamp, using the segmentation call that matches live input.
  5. In the listener, draw the frame and its mask together, as described in the next section.

Render the mask as a background effect

Rendering is where many prototypes go wrong, because a mask can look right while carrying the wrong values. Process each frame in this order:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Draw the current video frame to a canvas.
  2. Read the mask for that same frame. For a category mask, identify the person class from the label map of the model file you ship. Do not hard-code an index copied from another model or another tutorial.
  3. Compare the mask dimensions with the frame. If they differ, resize the mask explicitly rather than letting the browser stretch it.
  4. Composite the result: draw the blurred or replacement background first, then draw the person pixels on top, using the mask as alpha.
  5. If you mirror the preview, mirror the mask the same way. A mismatch shows up as a cutout that drifts relative to the person.

Each format conversion adds per-frame cost, so keep the mask in the representation you received for as long as the pipeline allows.

Fidelity limits and scope

The Selfie Segmentation model card, dated 6 May 2021 and written by Tingbo Hou, Siargey Pisarchyk, and Karthik Raveendran of Google, states: “The model is optimized for real-time performance in the web browser and on a wide variety of mobile devices, and may not provide pixel perfect masks.” Plan for that. The card names these conditions as sources of error:

  • Thin features such as fingers may occasionally be missed.
  • Poor lighting, image noise, fast motion, or a large occluding object can degrade mask quality.
  • The model may include multiple people of similar scale. People at different scales, and people farther than 14 feet (4 meters), are out of scope.

In practice, hair edges and fingers are where a background effect will show its errors first. The card also excludes surveillance and identity recognition and says the model is not intended for life-critical decisions. A blur for video calls matches that design; an analytics feature that depends on exact edges or identity does not. The official documentation gives no named accuracy percentage, so do not quote one.

Why segmentation breaks on Firefox or Safari

The three reports below are documented failures relevant to this stack. Each is tied to a specific version and environment. None shows that an entire browser is broken, and none proves that a current release still fails. Test the exact versions your users run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Firefox 135.0.1 with the GPU delegate

A MediaPipe issue opened 3 March 2025 reports that Image Segmenter fails with the GPU delegate on Firefox 135.0.1 using MediaPipe 0.10.9. The reporter also saw a WebGL readPixels warning about a format/type incompatibility. When checked, the issue was marked as awaiting a response from a Google engineer. Check the issue for updates before drawing conclusions about current Firefox releases.

iOS Safari: wrong category labels from the GPU delegate

A separate issue titled “[Web] Image Segmenter GPU delegate produces incorrect categories on iOS Safari” describes scrambled category labels from the GPU delegate. The reproduction used @mediapipe/tasks-vision 0.10.22-rc from March 2025, and the reporter said the CPU output was correct in that setup. This is the most dangerous failure in the list because nothing throws: a mask still appears, but its class IDs are wrong. Verify category distributions and visual overlays on each supported device, not just that a mask appears.

Legacy @mediapipe/selfie_segmentation under a strict CSP

A 2021 issue, opened 19 November 2021, says the legacy @mediapipe/selfie_segmentation JavaScript bindings failed under a Content Security Policy that disallowed unsafe-eval, with Chrome 96 and MediaPipe v0.8.5. The report traces the failure to dynamically generated code in that package. Treat it as a historical compatibility test. It does not show that the current @mediapipe/tasks-vision package has the same requirement, so test the exact package, bundler output, CSP, and browser your application deploys.

A debugging sequence for failures across browsers and devices

Record the following for every failure before changing code. Without these fields a report cannot be reproduced or compared:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Browser and exact version, operating system, device, and GPU
  • Package version (for example @mediapipe/tasks-vision) and the model asset loaded
  • Running mode (IMAGE, VIDEO, or LIVE_STREAM) and delegate (GPU or CPU)

Then work through the failure in this order:

  1. Compare delegates on the same frame. Run one captured frame through the GPU and CPU delegates. Compare category counts or mask pixels, then the composited result. A plausible-looking overlay can still carry wrong class IDs.
  2. Split still images from live input. If a still image is correct and the camera stream is not, suspect capture, orientation, or frame timing before suspecting the model.
  3. Log the failure points. Record model load errors, WebGL or WASM errors, and the time between sending a frame and receiving its result in the listener. Growing gaps between those timestamps point to lag.
  4. Keep the page working when a path fails. Show a recoverable state, and fall back to the CPU delegate where its measured speed is acceptable for your workload. Check the latency table first, because the CPU path can be far slower for some models.
  5. Test edge conditions. Cover hair and fingers, motion, dim light, noise, partial occlusion, and people at several distances.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Privacy and network behaviour

“On-device” describes where image processing happens. It does not mean the page makes no network requests. Google’s MediaPipe APIs terms, last updated 28 May 2026 (UTC), state: “When you use MediaPipe Solution APIs, processing of the input data (e.g. images, video, text) fully happens on-device, and MediaPipe does not send that input data to Google servers.” The same terms say the APIs may contact Google for bug fixes, updated models, and accelerator compatibility information. They also say the APIs may send performance and utilisation metrics, including inference counts, hardware-level performance, application and input metadata, and system environment. The terms place responsibility on the developer to obtain informed consent for metrics processing where that is required.

Your page also fetches the model file and the WebAssembly runtime from wherever you host them, so network traffic exists regardless of where inference runs. The accurate claim for your product is:

MediaPipe processes the input on-device, but the API can still contact Google and send usage or environment metrics.

Do not describe a page as making no network requests until you have checked asset delivery, your own telemetry, and every other dependency.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.