How can a camera determine that a person has opened their mouth?
At first, this sounds like a computer-vision problem. But underneath it is a much simpler mathematical idea:
facial observation → geometric measurement → numerical decision
A face-detection system can provide the first part: coordinates describing parts of the face. Mathematics can then turn those coordinates into a measurable quantity. Finally, an application can compare that quantity with a threshold and make a decision.
This simple chain is useful for understanding how an observable facial action can become one signal in a broader face liveness or Presentation Attack Detection (PAD) workflow.
It is also important to understand what the measurement does not prove.
From a Face to Measurable Geometry
A camera captures an image, but a mathematical calculation needs numbers.
Google ML Kit's Face Detection API can provide facial contours as collections of two-dimensional points. Among the available contours are UPPER_LIP_BOTTOM, representing the bottom outline of the upper lip, and LOWER_LIP_TOP, representing the top outline of the lower lip. [4]

Conceptually, the process looks like this:

image → facial contours → coordinates → mathematical measurement
The important transition is that a visual feature has become numerical data.
In the implementation examined here, one point is selected from each lip contour:
List<PointF> upperLipBottomContour =
face.getContour(FaceContour.UPPER_LIP_BOTTOM).getPoints();
List<PointF> lowerLipTopContour =
face.getContour(FaceContour.LOWER_LIP_TOP).getPoints();
The implementation then selects index 4 from each collection:

0 1 2 3 4 5 6 7 8
↑
index 4Because Java uses zero-based indexing, index 4 is the fifth point.
This detail matters because the selected point should not be described as a formally defined "mouth-center" landmark. Choosing index 4 is an implementation choice. ML Kit provides the contour points, while the application decides which points to use.
Let the selected upper-lip point be:
Uₜ = (xᵤ,ₜ, yᵤ,ₜ)
and the selected lower-lip point be:
Lₜ = (xₗ,ₜ, yₗ,ₜ)
The subscript t represents the current frame or observation.
We now have two points in a two-dimensional coordinate system.
Turning Two Points Into One Number
Once the two points are available, the next question is straightforward:
How far apart are they?
For two points in two-dimensional space, the Euclidean distance is:
dₜ = √((xₗ,ₜ − xᵤ,ₜ)² + (yₗ,ₜ − yᵤ,ₜ)²)
This is the familiar distance formula from analytic geometry.
The implementation calculates the same quantity:
double distance = Math.sqrt(
(x2 - x1) * (x2 - x1) +
(y2 - y1) * (y2 - y1)
);The result, dₜ, is a single numerical signal representing the separation between the two selected lip points.
When the mouth is relatively closed, the points should generally be closer together. When the mouth opens, their separation should generally increase.
That gives us a simple model:
facial contours → two points → one distance
A visual configuration has been reduced to one number.
A small numerical example
Suppose the detector produces:
Uₜ = (120, 150)
and:
Lₜ = (122, 185)
Then:
dₜ = √((122 − 120)² + (185 − 150)²)
= √(2² + 35²)
= √1229
≈ 35.06
The implementation discussed here uses a threshold of 30. Therefore:
35.06 > 30
and this observation satisfies the implementation's open-mouth condition.
The calculation itself is simple. The assumptions surrounding it are not.
From a Number to a Decision
A measurement does not automatically tell an application what to do.
The application needs a decision rule:
double openMouth = mouthOpenScore(face);
boolean isOpen = openMouth > 30.0;Mathematically:
OpenMouthₜ =
true if dₜ > 30
false if dₜ ≤ 30The complete idea is:

two points → distance → threshold → decision
This is the central mathematical model of the article.
But there is an important distinction between the two final steps.
The distance equation is mathematically well-defined:
dₜ = √((xₗ,ₜ − xᵤ,ₜ)² + (yₗ,ₜ − yᵤ,ₜ)²)
The threshold:
dₜ > 30
is an empirical engineering choice.
The value 30 is specific to the implementation being discussed. It is not a universal open-mouth threshold and should not be interpreted as a recommended security threshold.
A different camera, image resolution, detector configuration, face scale, or capture environment could produce different values.
This is a useful example of a broader principle in computational research:
A mathematical measurement can be precise even when the decision built on that measurement requires empirical validation.
Where Liveness Enters the Picture
An open-mouth measurement can be useful in an interactive workflow.
A system might ask a user to:
smile,
blink,
open their mouth,
or move their head.
The computer-vision system observes the requested action, converts part of that observation into a numerical signal, and determines whether the expected condition has been reached.

But an observable action is not the same thing as proof of liveness.
Presentation Attack Detection (PAD) addresses whether a biometric presentation is bona fide or represents an attack against the biometric capture process. ISO/IEC 30107-3 provides a framework for testing and reporting PAD performance. [1]
A recorded video, for example, may already contain the requested mouth movement. A predictable challenge could therefore be reproduced without demonstrating that a live person is responding to the challenge at that moment.
The distinction is:
open-mouth detection: "Do these facial coordinates satisfy the expected geometric condition?"
liveness/PAD: "Does the observed presentation provide sufficient evidence that it is a bona fide presentation rather than an attack?"
identity verification: "Does the observed face correspond to the claimed identity?"
These are related security functions, but they answer different questions.
The geometric measurement belongs to the first category. It can become one signal within the second, but it does not independently establish the second or third.
Why the Simple Measurement Has Limits
The model is attractive because it is easy to understand:
Uₜ, Lₜ → dₜ → dₜ > 30
But simplifying a visual phenomenon also introduces assumptions.
Face scale
The distance is measured in image coordinates rather than physical units.
A face that appears larger in the image can produce a larger distance even when the underlying facial movement is similar.
One possible improvement is normalization:
rₜ = dₜ / sₜ
where sₜ could represent a characteristic facial scale, such as face width.
The implementation discussed here does not perform this normalization.
Head pose
Turning or tilting the head changes how the facial geometry appears in a two-dimensional image. Yaw and pitch can therefore affect the measured separation between the selected points.
Image quality
Lighting, motion blur, low resolution, occlusion, and camera noise can affect the quality of the detected contours.
The Euclidean formula cannot correct inaccurate input coordinates. The mathematics can be perfectly correct while the measurement itself is unreliable.
Individual variation
People have different lip shapes, facial proportions, and natural ranges of mouth movement. A single fixed threshold can therefore behave differently across users.
These factors can result in false positives and false negatives in the open-mouth classification.
This is why the threshold cannot simply be assumed to work universally. It needs to be evaluated using representative data and appropriate experimental conditions.
What the Measurement Does Not Establish
The limitations become more important when the measurement is considered from a security perspective.
The method does not establish:
a universal open-mouth threshold;
a population-level accuracy rate;
resistance to specific presentation attack instruments;
measured APCER or BPCER;
superiority over another liveness method; or
production readiness.
Those claims would require a controlled evaluation with defined bona fide presentations, attack presentations, capture conditions, threshold-selection procedures, and appropriate PAD metrics.
For example, PAD research commonly uses Attack Presentation Classification Error Rate (APCER) and Bona Fide Presentation Classification Error Rate (BPCER) when evaluating performance. [1][3]
This article does not report those metrics because it does not conduct a controlled PAD experiment.
The security limitation can therefore be expressed simply:
An observable facial action can provide useful information within an interactive workflow, but observing that action does not by itself establish that the presentation comes from a live person or that the person's identity has been verified.
That distinction is important when translating a mathematical model into a security claim.
A Simple Rule in a Larger PAD Landscape
The open-mouth method represents one of the simplest ways to convert facial information into a decision.
More sophisticated face anti-spoofing systems can use substantially more information.
Earlier approaches often relied on handcrafted features describing properties such as texture, appearance, or motion. Modern approaches increasingly use learned representations from deep neural networks. Other systems incorporate temporal information, depth, infrared, or multiple sensing modalities. [2][3][6]
The difference is fundamentally about what evidence is being observed.
Approach | Main evidence | General characteristic |
|---|---|---|
Geometric rule | Facial coordinates | Simple and interpretable |
Handcrafted features | Texture and appearance | Uses richer manually designed features |
Deep learning | Learned image representations | Can model complex visual patterns |
Temporal models | Changes across frames | Uses motion and temporal context |
Depth / IR / multimodal | Additional sensing | Provides information beyond a single RGB image |
The geometric approach is therefore not intended to replace modern PAD systems.
Its value is that the entire reasoning chain can be inspected:
observation → measurement → decision
That makes it useful as a teaching example, an interpretable component, or a lightweight signal in a larger interaction.
More complex models can provide richer evidence, but they also introduce their own requirements, including training data, computational resources, and the challenge of generalizing to previously unseen attacks. [3][6]
The Broader Lesson: Measurement Is an Abstraction
The interesting part of this example is not really the mouth.
It is the abstraction.
A camera observes something visually complex: a human face changing shape.
A computer-vision system represents part of that face as coordinates.
Mathematics reduces those coordinates to a distance.
An engineering rule converts that distance into a Boolean decision.
Two points → one distance → one decision
That process is powerful because it makes a complicated observation computationally manageable.
It is also where assumptions enter the system.
The two points do not capture the entire face. The distance does not represent physical mouth size. The threshold does not have universal validity. And the resulting Boolean value does not prove liveness or identity.
This is a broader lesson that extends beyond biometric systems:
Simplifying a complex phenomenon into measurable variables makes computation possible, but the abstraction also determines what information is lost and what assumptions the final decision depends on.
That is where mathematics and security engineering meet.
Conclusion
A camera does not directly produce a statement such as "the user opened their mouth."
It produces an image.
A computer-vision system can turn part of that image into facial contours. From those contours, an application can select representative points. Geometry can then turn the points into a numerical measurement.
For the implementation examined here:
Uₜ, Lₜ → dₜ → dₜ > 30
where:
dₜ = √((xₗ,ₜ − xᵤ,ₜ)² + (yₗ,ₜ − yᵤ,ₜ)²)
The mathematics is straightforward.
The engineering interpretation requires more care.
Index 4 is an implementation choice, not a formally defined ML Kit mouth-center landmark. The value 30 is an implementation-specific threshold, not a universal security value. Image scale, head pose, image quality, and individual variation can all affect the measurement.
Most importantly, detecting an open mouth does not prove that a live person is in front of the camera. It is an observable facial signal that can contribute to a broader liveness or PAD workflow.
The value of the example is therefore not that two points can solve face liveness.
It is that two points demonstrate a much broader computational pattern:
a visual observation becomes geometry, geometry becomes a number, and a number becomes a decision.
Understanding what that decision actually measures, what assumptions it makes, and what it cannot prove is where the real security reasoning begins.
References
[1] ISO/IEC 30107-3:2023
International Organization for Standardization. ISO/IEC 30107-3:2023 — Information technology — Biometric presentation attack detection — Part 3: Testing and reporting. 2023.
This standard establishes principles and methods for assessing PAD performance, reporting testing results, and classifying known presentation attack types.
https://www.iso.org/standard/79520.html
[2] Sharma and Selwal — Face Presentation Attack Detection Survey
D. Sharma and A. Selwal. “A survey on face presentation attack detection mechanisms: hitherto and future perspectives.” Multimedia Systems, vol. 29, 2023, pp. 1527–1577.
DOI: 10.1007/s00530-023-01070-5
https://doi.org/10.1007/s00530-023-01070-5
[3] Yu et al. — Deep Learning for Face Anti-Spoofing
Z. Yu, Y. Qin, X. Li, C. Zhao, Z. Lei, and G. Zhao. “Deep Learning for Face Anti-Spoofing: A Survey.” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 5, 2023, pp. 5609–5631.
DOI: 10.1109/TPAMI.2022.3215850
https://doi.org/10.1109/TPAMI.2022.3215850
[4] Google ML Kit — Face Detection
Google. ML Kit — Face Detection. Documentation covering face detection, facial landmarks, contours, and related configuration.
https://developers.google.com/ml-kit/vision/face-detection
[5] Kredibel — VisionSample-Android
Kredibel. VisionSample-Android: An example of implementing Liveness Detection and Identity OCR on Android application using Kredibel-Vision-SDK.
The public sample demonstrates liveness-related detection flows including smile, mouth opening, blinking, and head movements.
https://github.com/kredibel-id/VisionSample-Android
[6] Xing et al. — Face Anti-Spoofing Based on Deep Learning
H. Xing, S. Y. Tan, F. Qamar, and Y. Jiao. “Face Anti-Spoofing Based on Deep Learning: A Comprehensive Survey.” Applied Sciences, vol. 15, no. 12, 2025, 6891.
DOI: 10.3390/app15126891
https://doi.org/10.3390/app15126891
[7] Zhang, Tondi, and Barni — Replay Attacks Against CNN-Based Anti-Spoofing
B. Zhang, B. Tondi, and M. Barni. “Adversarial examples for replay attacks against CNN-based face recognition with anti-spoofing capability.” Computer Vision and Image Understanding, vol. 197–198, 2020, 102988.
DOI: 10.1016/j.cviu.2020.102988
https://doi.org/10.1016/j.cviu.2020.102988
Disclaimer: The views, information, interpretations, and conclusions presented in this article are those of the author, who remains responsible for the accuracy, originality, and appropriate sourcing of the content. The Research Code provides an editorial and publishing platform and does not necessarily endorse the views or claims expressed in the article. Any matters arising from the content of the article remain the responsibility of the author.






