System and method of predicting human interaction with vehicles

Patent No. US10614344 (titled "System and method of predicting human interaction with vehicles") on Jul 16, 2019. The application was issued on Apr 7, 2020.

What is this patent about?

’344 is related to the field of autonomous vehicle navigation and predictive data analytics. Specifically, it addresses the technical challenge of enabling self-driving systems to anticipate the behavior of other road participants, such as pedestrians and cyclists, by mimicking the intuitive social reasoning and hazard perception of human drivers.

The underlying idea behind ’344 is that human intuition regarding road behavior can be quantified and transferred to a machine through a specialized training pipeline. Instead of relying solely on physical motion vectors, the system captures the collective subjective judgment of human observers who view road scenes and predicts the statistical distribution of their responses to estimate the likely intentions or awareness of others.

The claims of ’344 focus on a computer-implemented method that receives real-time sensor data from an autonomous vehicle and processes it through a supervised learning based model. This model is uniquely configured to output a statistical summary that characterizes how a population of human users would likely respond to that specific visual stimulus, using this aggregate human-like prediction to control the vehicle’s physical operation.

In practice, the invention works by presenting human observers with derived stimuli—manipulated video segments that highlight specific actors, such as a cyclist in a bounding box. The system records not just the observers' explicit predictions about an action, but also implicit data like response time and eye movement, aggregating these into a target dataset that represents a human consensus on road social dynamics.

This approach differs from prior solutions by moving beyond simple trajectory extrapolation. By training a neural network to predict the aggregate human perception of a road user's state of mind—such as whether a pedestrian is distracted or aggressive—the vehicle can make nuanced navigation decisions that reflect the complex, non-linear social interactions inherent in urban driving environments.

How does this patent fit in bigger picture?

Technical Landscape

In the late 2010s when ’344 was filed, autonomous vehicle systems and driver-assistance technologies were increasingly prevalent at a time when object detection and path planning were typically implemented using kinematic motion vectors. When systems commonly relied on the extrapolation of current and past physical movements to predict future trajectories, software constraints made the interpretation of human intent and social cues non-trivial. During this era, standard architectures focused on calculating the physical velocity and heading of pedestrians or cyclists rather than modeling the underlying psychological states or behavioral likelihoods that drive those movements.

Prosecution Position

The disclosed invention represents a technical advancement by shifting from purely kinematic trajectory extrapolation to a model-based prediction of human intent derived from aggregated human cognitive responses. The architectural solution involves a multi-stage pipeline where road scene stimuli are presented to human observers to collect response data—including both explicit action predictions and implicit metrics like response latency and eye-tracking—which are then aggregated into statistical distributions. This integration enables the training of supervised learning models, such as deep convolutional or recurrent neural networks, to output predicted summary statistics for live sensor data. The technical effect achieved is the ability of a vehicle to anticipate complex social behaviors, such as a pedestrian’s willingness to cross or a cyclist’s awareness of the vehicle, overcoming the constraints of traditional motion-vector systems that fail to account for the state of mind of road users.

Claims

This patent contains 22 claims, with claims 1, 10, and 19 serving as the independent claims. The independent claims focus on a system and method for controlling an autonomous vehicle by using a supervised learning model to predict a statistical distribution of expected human responses to captured sensor data showing road objects. The dependent claims serve to specify various machine learning architectures, define the types of road objects and user response parameters, and detail specific human factors being predicted, such as a road user's state of mind, intentions, or awareness of the vehicle.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Distribution of user responses
(Claim 1, Claim 10, Claim 19)
Summary statistics are generated based on the responses of all of the observers who looked at an image. Individual variability in responses to a given stimulus can be characterized in the information given by the observers to the learning algorithm. Because these summary statistics characterize the distribution of human responses that predict the state of mind of a road user, the predicted statistics are a prediction of the aggregate judgment of human observers.The range and frequency of different human judgments concerning a specific road scenario, used to capture individual variability in human prediction.
Sensor data
(Claim 1, Claim 10, Claim 19)
The 'real world' or 'live data' video or other sensor frames from a car-mounted sensor are delivered to the trained learning algorithm. In one embodiment, the data from the frame that was passed through the model would comprise the pixel data from a camera. These frames have the same resolution, color depth and file format as the frames used to train the algorithm.Raw information captured by vehicle-mounted hardware, such as cameras, representing a road scene and the objects within it.
Statistical summary
(Claim 1, Claim 10, Claim 19)
Summary statistics are generated based on the user responses and may characterize the aggregate responses of multiple human observers to a particular derived stimulus. These summary statistics could include measurements of the central tendency of the distribution of scores like the mean, median, or mode. They could also include measurements of the heterogeneity of the scores like variance, standard deviation, skew, kurtosis, heteroskedasticity, multimodality, or uniformness.Data characterizing the aggregate responses of multiple human observers to a road scene stimulus, including measures of central tendency or heterogeneity regarding predicted object behavior.
Supervised learning based model
(Claim 1, Claim 10, Claim 19)
The collection of images and statistics can be used to train a supervised learning algorithm, which can comprise a random forest regressor, a support vector regressor, a simple neural network, a deep convolutional neural network, a recurrent neural network, or a long-short-term memory (LSTM) neural network. The algorithm adapts its architecture in terms of weights or structure to minimize the deviation between its predicted label on a novel stimulus and the actual label collected on that stimulus. The model is optimized by progressively adjusting parameters in response to the characteristics of the images and summary statistics given to it in the training phase.An algorithm, such as a neural network or regressor, trained to minimize the deviation between its predictions and labels derived from human observer responses to road scene stimuli.
User responses
(Claim 1, Claim 10, Claim 19)
The response can be categorized in terms of how many human observers believe that the pedestrian will stop upon reaching the intersection, continue walking straight, or turn a corner. The explicit response of the observer is recorded as well as implicit data, which can include how long the subject took to respond, if they hesitated, if they deleted keystrokes, or where their eyes moved. These responses characterize the distribution of human judgments that predict the state of mind of a road user.Explicit and implicit data provided by human observers regarding the predicted actions, intentions, or states of mind of road users pictured in a stimulus.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
2:25-cv-00742Jul 23, 2025Perceptive Automata Llc V. Tesla, Inc.

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10614344

Application Number
US16512560A
Filing Date
Jul 16, 2019
Publication Date
Apr 7, 2020
External Links
Slate, USPTO , Google Patents