Visual recognition using user tap locations

Patent No. US10353950 (titled "Visual recognition using user tap locations") on Jun 28, 2016. The application was issued on Jul 16, 2019.

What is this patent about?

’950 is related to the field of visual search engines and image recognition. Specifically, it addresses the challenge of disambiguating user intent when a query image contains multiple objects or complex backgrounds. Traditional visual search often struggles to identify which specific element within a photograph a user actually wants to learn about, leading to irrelevant results and wasted computational resources.

The underlying idea behind ’950 is to use a coordinate-based user interaction, such as a screen tap, as a spatial filter to prioritize image processing. By combining a user tap location with automated object detection, the system can bypass the ambiguity of a full-image search. The core insight is to treat the tap not just as a pointer, but as a trigger to isolate a specific entity within a bounding box, thereby focusing the search engine's attention on the user's precise area of interest.

The claims of ’950 focus on a method that receives both a query image and a tap location to drive a targeted search workflow. The process utilizes an object detection neural network to map out various entities in the image using bounding boxes. The system then determines which specific entity’s bounding box contains the tap coordinates. Once the entity is isolated, the system retrieves and presents a set of pre-associated textual search queries specifically relevant to that object, rather than the image as a whole.

In practice, the invention functions by dynamically linking spatial data to a knowledge graph. When a user selects a point on their device, the system calculates the proximity of that point to detected objects. If the tap falls within a defined bounding box, the system generates a list of suggested queries—such as specific questions about a building's height or an athlete's statistics—which are then displayed as actionable options. This allows the user to transition from a visual input to a refined textual search with a single gesture.

This approach differs from prior solutions by moving away from global image descriptors that often produce generic results. Instead of attempting to classify the entire scene, ’950 uses the tap to perform content-aware filtering, effectively ignoring irrelevant background noise. By narrowing the scope of the search to a specific identified entity before the query is even executed, the system reduces the cognitive load on the user and ensures that the final search results are highly aligned with the user's immediate intent.

How does this patent fit in bigger picture?

Technical Landscape

In the mid-2010s when ’950 was filed, visual search systems were typically implemented using server-side recognition engines that processed entire image frames to identify objects or text. At a time when mobile hardware constraints and network latency made the exhaustive processing of high-resolution images non-trivial, systems commonly relied on global image descriptors or uniform feature extraction across the whole field of view rather than localized user-guided analysis. Consequently, computational resources were often distributed equally across relevant and irrelevant image regions, which could lead to high latency or the identification of background entities that did not align with the user's specific intent.

Prosecution Position

The disclosed invention represents a meaningful technical advancement by integrating user-defined spatial coordinates—specifically a tap location—directly into the visual recognition pipeline to bias computational resource allocation. This architectural shift moves away from uniform image processing toward a targeted approach where the system performs content-aware cropping, higher-density descriptor extraction, or more intensive neural network classification specifically within an area of interest defined by the user. The technical effect achieved is a dual optimization: it reduces the total computational load by applying less expensive processing to background regions while simultaneously increasing recognition accuracy for the user's intended subject. Furthermore, the system enables improved query relevance by using the broader image context to disambiguate entities identified at the specific tap location, overcoming the constraint of isolated object recognition.

Claims

This patent contains 17 claims, with claims 1, 16, and 17 serving as the independent claims. The independent claims focus on a computer-implemented process for identifying specific entities within a query image by correlating a user's tap location with bounding boxes generated by an object detection neural network, subsequently presenting pre-associated textual search queries and results based on that proximity. The dependent claims serve to refine this process by detailing specific image processing techniques such as content-aware cropping, the use of classification neural networks and optical character recognition, descriptor matching within areas of interest, and methods for scoring the relevance of suggested textual queries.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Bounding box
(Claim 1, Claim 16, Claim 17)
The knowledge engine may define a bounding box around each identified one or more entities that are associated with a processed query image. The knowledge engine may then determine that the user tap location lies within one or more bounding boxes of one or more respective entities and assign a higher relevance score to the one or more respective entities than other identified entities.A defined geometric boundary or coordinate range that encapsulates a detected entity within a query image, used to determine if a user's selection overlaps with that entity.
Object detection neural network
(Claim 1, Claim 16, Claim 17)
Image recognition systems and procedures can be computationally expensive, since effectively recognizing objects or text in images may involve processing an image using a deep neural network, e.g., a convolutional neural network. The system allows visual recognition engines to effectively apply visual recognition resources, such as neural networks, to areas of an image that a user is interested in. The system allocates and applies more processing power to an area of an image that a user has indicated as being important.A machine learning model, such as a convolutional neural network, used to identify and locate specific entities within an image by defining their spatial boundaries.
Pre-associated
(Claim 1, Claim 16, Claim 17)
The knowledge engine can identify information that is pre-associated with the one or more entities. The database or server may include a trained or hardcoded statistical mapping of entities, e.g., based on search query logs, and can store candidate search queries that relate to various entities. The knowledge engine can identify the related candidate search queries by accessing entries at the database or server that are distinctly related to the identified entity.A stored relationship or statistical mapping between a specific recognized entity and a set of relevant search queries or data established prior to the current search request.
Suggested textual search queries
(Claim 1, Claim 16, Claim 17)
The knowledge engine can identify information that is pre-associated with the one or more entities. For example, the knowledge engine can access the database or server to identify candidate search queries that are associated with the entity “The Gherkin,” such as “How tall is The Gherkin” or “Directions to The Gherkin” using a pre-computed query map. The database or server may include a trained or hardcoded statistical mapping of entities, e.g., based on search query logs, and can store candidate search queries that relate to various entities.A set of text-based queries pre-linked to recognized entities, presented to the user as selectable options to retrieve specific information related to the image content.
User tap location
(Claim 1, Claim 16, Claim 17)
A system can receive a query image and a user tap location, e.g., a photograph from a user's surroundings with a selected area of interest. The visual recognition results are improved by using the user tap location. For example, visual recognition results may be used to enhance inputs to backend recognizers and may be used to rank obtained recognition results.A specific coordinate or area within a query image selected by a user to indicate a particular object or region of interest for visual recognition.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:25-cv-00514Apr 29, 2025Accusearch Technologies Llc V. Google Llc

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10353950

Application Number
US15195369A
Filing Date
Jun 28, 2016
Publication Date
Jul 16, 2019
External Links
Slate, USPTO , Google Patents