Patent No. US10353950 (titled "Visual recognition using user tap locations") on Jun 28, 2016. The application was issued on Jul 16, 2019.
’950 is related to the field of visual search engines and image recognition. Specifically, it addresses the challenge of disambiguating user intent when a query image contains multiple objects or complex backgrounds. Traditional visual search often struggles to identify which specific element within a photograph a user actually wants to learn about, leading to irrelevant results and wasted computational resources.
The underlying idea behind ’950 is to use a coordinate-based user interaction, such as a screen tap, as a spatial filter to prioritize image processing. By combining a user tap location with automated object detection, the system can bypass the ambiguity of a full-image search. The core insight is to treat the tap not just as a pointer, but as a trigger to isolate a specific entity within a bounding box, thereby focusing the search engine's attention on the user's precise area of interest.
The claims of ’950 focus on a method that receives both a query image and a tap location to drive a targeted search workflow. The process utilizes an object detection neural network to map out various entities in the image using bounding boxes. The system then determines which specific entity’s bounding box contains the tap coordinates. Once the entity is isolated, the system retrieves and presents a set of pre-associated textual search queries specifically relevant to that object, rather than the image as a whole.
In practice, the invention functions by dynamically linking spatial data to a knowledge graph. When a user selects a point on their device, the system calculates the proximity of that point to detected objects. If the tap falls within a defined bounding box, the system generates a list of suggested queries—such as specific questions about a building's height or an athlete's statistics—which are then displayed as actionable options. This allows the user to transition from a visual input to a refined textual search with a single gesture.
This approach differs from prior solutions by moving away from global image descriptors that often produce generic results. Instead of attempting to classify the entire scene, ’950 uses the tap to perform content-aware filtering, effectively ignoring irrelevant background noise. By narrowing the scope of the search to a specific identified entity before the query is even executed, the system reduces the cognitive load on the user and ensures that the final search results are highly aligned with the user's immediate intent.
In the mid-2010s when ’950 was filed, visual search systems were typically implemented using server-side recognition engines that processed entire image frames to identify objects or text. At a time when mobile hardware constraints and network latency made the exhaustive processing of high-resolution images non-trivial, systems commonly relied on global image descriptors or uniform feature extraction across the whole field of view rather than localized user-guided analysis. Consequently, computational resources were often distributed equally across relevant and irrelevant image regions, which could lead to high latency or the identification of background entities that did not align with the user's specific intent.
The disclosed invention represents a meaningful technical advancement by integrating user-defined spatial coordinates—specifically a tap location—directly into the visual recognition pipeline to bias computational resource allocation. This architectural shift moves away from uniform image processing toward a targeted approach where the system performs content-aware cropping, higher-density descriptor extraction, or more intensive neural network classification specifically within an area of interest defined by the user. The technical effect achieved is a dual optimization: it reduces the total computational load by applying less expensive processing to background regions while simultaneously increasing recognition accuracy for the user's intended subject. Furthermore, the system enables improved query relevance by using the broader image context to disambiguate entities identified at the specific tap location, overcoming the constraint of isolated object recognition.
This patent contains 17 claims, with claims 1, 16, and 17 serving as the independent claims. The independent claims focus on a computer-implemented process for identifying specific entities within a query image by correlating a user's tap location with bounding boxes generated by an object detection neural network, subsequently presenting pre-associated textual search queries and results based on that proximity. The dependent claims serve to refine this process by detailing specific image processing techniques such as content-aware cropping, the use of classification neural networks and optical character recognition, descriptor matching within areas of interest, and methods for scoring the relevance of suggested textual queries.
Definitions of key terms used in the patent claims.
US Latest litigation cases involving this patent.

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.
Get instant alerts for new documents