Patent No. US10614344 (titled "System and method of predicting human interaction with vehicles") on Jul 16, 2019. The application was issued on Apr 7, 2020.
’344 is related to the field of autonomous vehicle navigation and predictive data analytics. Specifically, it addresses the technical challenge of enabling self-driving systems to anticipate the behavior of other road participants, such as pedestrians and cyclists, by mimicking the intuitive social reasoning and hazard perception of human drivers.
The underlying idea behind ’344 is that human intuition regarding road behavior can be quantified and transferred to a machine through a specialized training pipeline. Instead of relying solely on physical motion vectors, the system captures the collective subjective judgment of human observers who view road scenes and predicts the statistical distribution of their responses to estimate the likely intentions or awareness of others.
The claims of ’344 focus on a computer-implemented method that receives real-time sensor data from an autonomous vehicle and processes it through a supervised learning based model. This model is uniquely configured to output a statistical summary that characterizes how a population of human users would likely respond to that specific visual stimulus, using this aggregate human-like prediction to control the vehicle’s physical operation.
In practice, the invention works by presenting human observers with derived stimuli—manipulated video segments that highlight specific actors, such as a cyclist in a bounding box. The system records not just the observers' explicit predictions about an action, but also implicit data like response time and eye movement, aggregating these into a target dataset that represents a human consensus on road social dynamics.
This approach differs from prior solutions by moving beyond simple trajectory extrapolation. By training a neural network to predict the aggregate human perception of a road user's state of mind—such as whether a pedestrian is distracted or aggressive—the vehicle can make nuanced navigation decisions that reflect the complex, non-linear social interactions inherent in urban driving environments.
In the late 2010s when ’344 was filed, autonomous vehicle systems and driver-assistance technologies were increasingly prevalent at a time when object detection and path planning were typically implemented using kinematic motion vectors. When systems commonly relied on the extrapolation of current and past physical movements to predict future trajectories, software constraints made the interpretation of human intent and social cues non-trivial. During this era, standard architectures focused on calculating the physical velocity and heading of pedestrians or cyclists rather than modeling the underlying psychological states or behavioral likelihoods that drive those movements.
The disclosed invention represents a technical advancement by shifting from purely kinematic trajectory extrapolation to a model-based prediction of human intent derived from aggregated human cognitive responses. The architectural solution involves a multi-stage pipeline where road scene stimuli are presented to human observers to collect response data—including both explicit action predictions and implicit metrics like response latency and eye-tracking—which are then aggregated into statistical distributions. This integration enables the training of supervised learning models, such as deep convolutional or recurrent neural networks, to output predicted summary statistics for live sensor data. The technical effect achieved is the ability of a vehicle to anticipate complex social behaviors, such as a pedestrian’s willingness to cross or a cyclist’s awareness of the vehicle, overcoming the constraints of traditional motion-vector systems that fail to account for the state of mind of road users.
This patent contains 22 claims, with claims 1, 10, and 19 serving as the independent claims. The independent claims focus on a system and method for controlling an autonomous vehicle by using a supervised learning model to predict a statistical distribution of expected human responses to captured sensor data showing road objects. The dependent claims serve to specify various machine learning architectures, define the types of road objects and user response parameters, and detail specific human factors being predicted, such as a road user's state of mind, intentions, or awareness of the vehicle.
Definitions of key terms used in the patent claims.
US Latest litigation cases involving this patent.

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.
Get instant alerts for new documents