Adjusting video layout

Patent No. US9088694 (titled "Adjusting video layout") on Oct 3, 2013. The application was issued on Jul 21, 2015.

What is this patent about?

’694 is related to the field of multi-party video conferencing and the dynamic management of visual layouts. In traditional continuous presence systems, participants join using a wide variety of hardware, ranging from large conference room monitors to small mobile displays. This diversity often leads to inconsistent viewing experiences where a participant’s face may appear too small to see on a mobile device or uncomfortably large on a high-resolution monitor, primarily because the system lacks a mechanism to normalize the visual scale of participants across different viewing environments.

The underlying idea behind ’694 is to achieve visual uniformity by dynamically adjusting the composition of video feeds based on the detected size of a participant's face relative to the total frame. Rather than simply tiling raw video streams, the system treats the participant's face as a specific region of interest that can be isolated and scaled. By calculating the ratio of this region to the full image, the system can mathematically normalize the appearance of all participants, ensuring that everyone appears at a consistent relative size regardless of how far they sit from their respective cameras.

The claims of ’694 focus on a multi-conferencing unit (MCU) that performs automated image manipulation to standardize layouts. The independent claims describe a process of detecting a region of interest within incoming feeds, calculating the ratio between that region and the full image area, and then centering and cropping the resulting image. This ensures the participant remains the focal point of the view-port. Additionally, the claims cover a feedback mechanism where the MCU detects the screen sizes of all participants and signals a preferred field-of-view setting to the endpoints to optimize the source video for the smallest display in the group.

In practice, the invention functions as an intelligent director that sits between the callers. When a participant joins from a smartphone, the MCU recognizes the limited screen real estate and can instruct other endpoints to tighten their camera field of view or perform the crop itself at the server level. By selecting a scale based on the maximum ratio of face-to-frame area among all participants, the system eliminates the jarring transition between a wide-angle room view and a close-up webcam shot, creating a cohesive gallery view where every participant is framed similarly.

This approach differs from prior solutions that relied on manual pan-tilt-zoom controls or static layout templates that ignored the actual content of the video feed. While older systems might simply resize a window to fit a grid, this invention utilizes face detection metadata to actively re-compose the internal geometry of the video stream. By centering the region of interest and scaling the cropped result to a target resolution, the system provides a high-quality, aesthetically pleasing experience that adapts to the specific hardware constraints of every individual participant in the call.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’694 was filed, multi-party video conferencing was typically implemented using a centralized Multipoint Control Unit (MCU) to aggregate and distribute video streams in a Continuous Presence layout. At a time when systems commonly relied on fixed camera fields-of-view or manual pan-tilt-zoom adjustments, maintaining visual consistency across diverse participant endpoints was difficult. Hardware and software constraints made it non-trivial to normalize the appearance of participants who were positioned at varying distances from their cameras or using different screen sizes, often resulting in layouts where some faces appeared disproportionately small or large relative to their assigned view-ports.

Prosecution Position

The disclosed invention represents a technical advancement in video layout management through an architectural shift that automates the normalization of participant appearances within a multi-party conference. By integrating real-time region-of-interest detection with dynamic scaling and cropping at the MCU, the system overcomes the constraint of static or mismatched camera feeds. The solution calculates a ratio between detected regions of interest and full image areas across all participants to determine a uniform relative face size for a specific layout. This capability enables the system to center and scale individual feeds so that all participants appear with consistent proportions, regardless of their original camera framing or the display constraints of the receiving endpoints.

Claims

The patent contains a total of 12 claims, with claims 1, 7, and 10 serving as the independent claims. These independent claims focus on methods and systems for optimizing video conference layouts by detecting regions of interest, such as participant faces, and adjusting image ratios, centering, and screen size signaling to improve visual composition. The dependent claims serve to further define the technical implementation by specifying the hardware components responsible for detection, the use of metadata for transmitting region information, and specific cropping techniques to fit desired display sizes.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Inducing endpoints to trim a camera's field of view
(Claim 7)
The MCU signals to each participant the preferred screen size for the conference. This can induce an endpoint to change its cameras' field of view to allow for the other side's screen preference. This essentially makes the video more suitable for viewing on a smaller screen.The process of influencing remote camera settings or output frames by signaling a preferred screen size to the endpoint, causing the endpoint to adjust what is captured or sent.
Multi-conferencing unit
(Claim 10)
In an embodiment of the invention, a system and method is provided to implement automatic scaling by a media conferencing unit (MCU) which is responsible for mixing the video streams sent by participants. The MCU composes the video streams into unique layouts and video streams sent back to each participant for viewing.A central system component (MCU) responsible for receiving, processing, and mixing video streams from multiple participants to compose and distribute unique layouts.
Ratio between each region of interest and a full image area
(Claim 1, Claim 10)
For instance, “in_ratio(k)” (where k is a participant identifier) can be used to refer to the ratio between the ROI area and the full image area in participant k's input stream. The algorithm, whether performed by the MCU or the participant's camera, may utilize a ratio to obtain a desired output.A calculated value (referred to as 'in_ratio(k)') representing the proportion of the total input video frame occupied by the detected region of interest.
Region of interest
(Claim 1, Claim 10)
The ROI may be the participant's face. It is understood that face detection algorithms may be used to detect a participant's face. The face detection algorithm may be performed by the MCU as it decodes the incoming video stream or by a participant's camera and sent to the MCU as meta-data.A specific portion of a video feed, typically identified as the participant's face using detection algorithms, used as the focal point for layout adjustments.
Relative size of each participant's face
(Claim 1, Claim 10)
The MDU determines the relative size of the participant's face that is to be shown for each participant in that layout. This may be done by selecting a scale that is equal to the max “in-ration) for all participants that are to be seen in a layout 1. The MCU then crops the image to fit the desired relative size.A scaling factor determined for a specific layout, often based on the maximum ratio of face-to-image area among all participants to be displayed in that layout.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
7:25-cv-00483Oct 22, 2025Arlington Technologies LLC v. NVIDIA Corporation

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US9088694

Application Number
US14045616A
Filing Date
Oct 3, 2013
Publication Date
Jul 21, 2015
External Links
Slate, USPTO , Google Patents