Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Google’s Pixel 2 created portrait-style background blur by combining three kinds of computation: an HDR+ photograph, a neural-network foreground mask, and a depth map derived from the rear camera’s dual-pixel sensor. The mask answered “which pixels belong to the person?” while depth answered “how far away is each region?” A renderer then used both to produce synthetic, depth-aware defocus.

This is a documented Pixel 2 and Pixel 2 XL design announced in 2017—not a guarantee that every later Pixel uses the same model, sensor path, or processing sequence.

Semantic segmentation in plain English

Semantic segmentation is dense, pixel-level classification. Instead of assigning one label to an entire photograph or drawing a box around an object, a model predicts a class or foreground probability for every pixel. In a portrait application, the useful result may be a mask whose values indicate how likely each pixel is to belong to a person.

The mask is an estimate, not a mathematically perfect boundary. Camera software can smooth, refine, or soften it before compositing. Hair, glasses, fingers, transparent materials, and motion-blurred edges are especially difficult because a pixel may contain both subject and background.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Google Pixel 11 Pro - Unlocked Smartphone, Gemini - 256 GB - Obsidian
  • Attention-grabbing design meets the latest evolution of the Google Pixel Camera on the new Google Pixel 11 Pro; Gemini Intelligence helps manage details so you can live in the moment[1]; and the phone is available in two sizes
  • Unlocked Android phone gives you the flexibility to change carriers and choose your own data plan: Works with Google Fi, Verizon, T-Mobile, AT&T, and other major carriers[2]
  • Stay informed without looking at your screen: When your phone is face down, Pixel HiLight gently alerts you with subtle glowing lights when your favorite contacts are calling or you’re talking with Gemini; exclusive to Google Pixel 11 Pro phones
  • Magic Capture catches the moment as you live it: With just one tap, Pixel 11 Pro captures video and photos, and automatically edits, crops, and unblurs a curated collection, ready to share – and you get the memory of how it felt to be in the moment
  • Two new cameras for more brilliant photos: A larger telephoto sensor captures 30% more light for clear, beautiful photos and videos, even in the dark[3]; Pixel’s longest zoom ever helps you capture details from impressive distances[4]
Technique Main question Typical output
Image classification What is in the image? One or more image labels
Object detection Where are the objects? Bounding boxes and labels
Semantic segmentation Which class does each pixel belong to? A pixel-level class or probability mask
Instance segmentation Which pixels belong to each individual object? A separate mask for every instance
Depth estimation How far away is each pixel or region? A depth map, often relative rather than absolute
Matting What fraction of each pixel is foreground? A soft alpha or transparency mask

Google’s Pixel 2 explanation calls its person-separation step semantic segmentation. In practice, it was specialized for portrait subjects rather than being a generic street-scene model: Google described training examples involving people, hats, sunglasses, and objects such as ice-cream cones.

Why a phone needs segmentation for portrait blur

A normal phone photograph is generally sharp across much of the frame. To imitate shallow depth of field, software must decide which pixels to preserve, which to blur, and how strongly to blur each background region. A rectangular crop would cut through hair, arms, and clothing; a green-screen method would require a controlled background. A learned mask lets the camera separate a person from an arbitrary scene.

Segmentation alone is not enough. A binary mask could keep the person sharp but would tend to apply one blur treatment to everything outside the mask. It would not know that a nearby object is in front of the subject or that a distant background should be blurred more than an object near the focus plane. That is why Google combined semantic understanding with geometric depth.

The Pixel 2 Portrait Mode pipeline

Google’s publicly documented flow can be summarized as:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

HDR+ image → neural foreground mask → dual-pixel depth map → depth-aware synthetic defocus

Rank #2
Google Pixel 10a - 30+ Hours Battery, Camera Coach, Gemini - Obsidian 128GB
  • Google Pixel 10a is a durable, everyday phone with more[1]; snap brilliant photography on a simple, powerful camera, get 30+ hours out of a full charge[2], and do more with helpful AI like Gemini[3]
  • Unlocked Android phone gives you the flexibility to change carriers and choose your own data plan; it works with Google Fi, Verizon, T-Mobile, AT&T, and other major carriers
  • Pixel 10a is sleek and durable, with a super smooth finish, scratch-resistant Corning Gorilla Glass 7i display, and IP68 water and dust protection[4]
  • The Actua display with 3,000-nit peak brightness shows up clear as day, even in direct sunlight[5]
  • Plan, create, and get more done with help from Gemini, your built-in AI assistant[3]; have it screen spam calls while you focus[6]; chat with Gemini to brainstorm your meal plan[7], or bring your ideas to life with Nano Banana[8]

1. HDR+ produces the base photograph

Portrait Mode began with an HDR+ image. HDR+ captured a burst of underexposed frames, aligned and averaged them, and used the combined data to reduce noise and improve highlight and shadow detail. A cleaner base image gives the segmentation network and stereo calculation more usable visual information. The blur was rendered over that processed result, not over a simple single exposure. This description comes from Google’s Pixel 2 account and should not be assumed to describe every current Pixel pipeline. Google Research

2. A CNN predicts the foreground mask

Google said it trained a convolutional neural network with skip connections on nearly one million pictures of people. The network estimated which pixels belonged to people and ran on the phone using TensorFlow Mobile.

At a high level, early convolutional layers respond to edges, colors, and textures. Deeper layers recognize larger structures such as faces, limbs, clothing, and body arrangements. Skip connections carry fine spatial information from earlier layers into later layers, helping the network make a high-level decision without losing as much boundary detail.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Google did not name a complete production architecture in that Pixel 2 article. It therefore would be inaccurate to state as fact that the phone used a particular DeepLab, MobileNetV2, or other named model. DeepLab-style networks are relevant industry context, but a Qualcomm developer page associating DeepLabV3+ with mobile image segmentation is not proof of the Pixel 2 production architecture. Qualcomm’s developer example

3. Dual-pixel data supplies a stereo depth cue

The Pixel 2 rear camera did not need two physically separate rear cameras to obtain stereo information. Its PDAF, or dual-pixel, sensor could form slightly different views through opposite sides of the lens. Google said the viewpoints were separated by less than approximately 1 millimeter, yet that small parallax could provide useful depth under favorable conditions.

The documented process formed left and right views, aligned them with a stereo algorithm, generated a lower-resolution depth map, and interpolated or refined it to higher resolution. Burst frames helped reduce noise and improve the estimate. The tiny baseline is also the source of important limitations: low light, textureless surfaces, repeated patterns, and movement make correspondence harder.

4. Mask and depth are rendered together

The mask identifies likely subject pixels; the depth map estimates relative distance. The renderer can keep the person comparatively sharp, vary blur according to distance from the focus plane, and treat objects in front of the person differently from a far background.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Google described a synthetic approximation of optical defocus. Rather than applying the same Gaussian blur everywhere, the system composites pixels with variable-sized translucent disks in depth order, approximating the disk-shaped blur associated with lens bokeh. The result can look convincing, but it is not identical to defocus formed optically by a large-aperture camera.

Rear camera and front camera were not equivalent

Rear camera

  • HDR+ supplied the base image.
  • The neural network predicted a person mask.
  • Dual-pixel/PDAF information supplied stereo depth.
  • The renderer combined semantic and geometric information for the portrait effect.

Front camera

  • It also used machine-learning segmentation.
  • It lacked PDAF pixels, so it did not provide the same stereo depth cue.
  • Its blur could not vary through a corresponding rear-camera depth map in the same way.

This hardware difference is why “Portrait Mode” should not be treated as one single algorithm. The inputs available to the camera determine what the software can infer.

Segmentation is not professional alpha matting

Segmentation predicts a semantic label or probability. Matting estimates partial foreground coverage, which is particularly valuable for fine hair, fur, translucent fabric, smoke, and motion-blurred boundaries. A camera may soften or refine a segmentation mask, but that does not establish that Pixel 2 used a separate neural matting stage.

Rank #4
Sale
Google Pixel 10 Pro - Unlocked Smartphone with Gemini - Obsidian - 128 GB
  • Google Pixel 10 Pro is the ultimate Pixel experience, featuring advanced AI with Gemini, unbelievable camera quality, impeccable design in two sizes, and the next-gen Google Tensor G5 chip[1]
  • Unlocked Android phone gives you the flexibility to change carriers and choose your own data plan[2]; it works - Google Fi, Verizon, T-Mobile, AT&T, and other major carriers
  • Get a head start on syncing your data before it even arrives: After you purchase your new Pixel, look for an email that explains how to transfer your photos, videos, passwords, and more in just a few quick steps[11]
  • Pixel’s pro camera system makes everything look amazing, even in low light; capture more of the scene with advanced Google AI models, and bring out incredible details with 100x Pro Res Zoom, stunning 50 MP images, and super steady videos in 8K[10]
  • Pixel 10 Pro is built with durable aluminum and Corning Gorilla Glass Victus 2 for scratch and drop resistance; the 6.3-inch Super Actua display with 3,300-nit peak brightness is easy on the eyes, even in direct sunlight[3,13,18]

The distinction explains common artifacts: hair can be partly classified as background, blur can bleed over a subject edge, and high-contrast boundaries can develop bright or dark halos. Transparent and reflective objects are difficult because their appearance does not map cleanly to a single foreground class.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why deep learning can run on a phone

A mobile vision model must balance boundary quality against latency, memory, battery consumption, heat, and preview responsiveness. Designers may reduce resolution, quantize weights, use specialized accelerators, or choose a smaller architecture. They also have to decide whether a model should specialize in people or generalize to pets, food, sky, and other subjects.

Google’s MobileNet research illustrates the broader engineering trade-off. MobileNetV2 was designed for on-device vision, including semantic segmentation, and Google reported that it used fewer parameters and operations than MobileNetV1 and ran approximately 30–40% faster on a Google Pixel phone in the comparison it presented. That benchmark is context for mobile-model design, not evidence that MobileNetV2 was the Pixel 2 Portrait Mode network. Google Research on MobileNetV2

Google stated that the Pixel 2 segmentation inference ran on the phone using TensorFlow Mobile. That supports narrower conclusions—lower dependence on a network connection and potentially lower latency for this step—but it does not prove that every Pixel camera feature is always processed on-device or that no image data is ever handled elsewhere.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, close subjects, and non-human objects

Google said Pixel 2 Portrait Mode processed a photograph in approximately four seconds and worked automatically, unlike the earlier Lens Blur mode that required moving the phone vertically. This is a historical Pixel 2 statement, not a current benchmark.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Google Pixel 7-5G Android Phone - Unlocked Smartphone with Wide Angle Lens and 24-Hour Battery - 256GB - Lemongrass
  • Google Pixel 7 is powered by Google Tensor G2; it’s faster, more efficient, and more secure, with the best photo and video quality yet on Pixel[1].Other camera description:Front,Rear.Bluetooth Version 5.2 with dual antennas for enhanced quality and connection.
  • Unlocked Android 5G phone gives you the flexibility to change carriers and choose your own data plan[2]; works with Google Fi, Verizon, T-Mobile, AT&T, and other major carriers
  • Pixel’s Adaptive Battery can last over 24 hours; when Extreme Battery Saver is turned on, it can last up to 72 hours[3]
  • The 6.3-inch Pixel 7 display is super sharp, with rich, vivid colors; it’s fast and responsive for smoother gaming, scrolling, and moving between apps[4]
  • Google Pixel 7 has wide and ultrawide lenses with up to 8x Super Res Zoom[5]; and Cinematic Blur brings more drama to your videos

When aimed at a small object such as a flower or food, the person-segmentation network could not provide a useful person mask. Google said the system could still use the depth map alone for nearby objects, working best at roughly less than one meter, while the camera could not focus sharply closer than approximately 10 centimeters. This fallback demonstrates that portrait effects were not the same as generic object-aware blur.

Where the system breaks

Segmentation errors

  • Frizzy, backlit, or partially hidden hair
  • Floppy hats, scarves, and unusual poses
  • Objects held close to the body
  • Multiple people with overlapping silhouettes
  • Transparent, reflective, or unfamiliar objects

Depth errors

  • Noise in low light
  • Blank walls with few matching features
  • Plaid and other repeated patterns
  • Thin structures and nearly identical subject/background distances
  • Motion between burst frames

Google specifically mentioned blank walls, plaid or repeating textures, and strong horizontal or vertical patterns as difficult stereo cases. Errors in the HDR+ image, mask, or depth map can propagate into the final portrait. The visible result may include halos, blur leakage, flat-looking blur, incorrectly sharp background patches, or incorrect treatment of an object in front of the subject.

What the Pixel example says about computational photography

The Pixel 2 illustrates a software-defined camera: the lens and sensor provide measurements, while learned perception, multi-frame processing, stereo cues, and rendering turn those measurements into a photographic effect. The neural network does not “understand” a scene in a human sense; it predicts likely pixel ownership from patterns learned during training. Depth then adds information that semantic identity cannot provide.

Google’s 2017 documentation is a clear historical example of semantic segmentation in a consumer camera. Later Pixel generations may use different sensors, models, accelerators, and learned-depth techniques. Without separate first-party documentation, the Pixel 2 sequence should not be presented as the architecture of every newer phone.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.