Precision

What Is Precision?

Precision is an evaluation metric used in classification. It measures how many predicted positives were actually positive.

Formula

Precision = True Positives / All Predicted Positives

When the model predicted positive, how often was it correct?

Precision focuses on the quality of the model's positive predictions.

Why Precision Matters

Precision matters when false positives are costly or harmful. A false positive happens when the model predicts positive, but the true answer is negative.

Spam Detection Example
  • Prediction: Spam
  • True Label: Not Spam
  • Problem: An important email may be sent to the spam folder.

High precision is important because users do not want normal emails incorrectly marked as spam.

Predicted Positives

Predicted positives are all examples the model labeled as positive. They include both true positives and false positives.

Disease Model Example
  • Predicted Positive: 50 patients
  • Actually Positive: 40 patients
  • Actually Negative: 10 patients
  • True Positives: 40, False Positives: 10
  • Precision = 40 / 50 = 80%

True Positives

A true positive happens when the model predicts positive and the true label is also positive. True positives are correct positive predictions.

Disease Detection Example
  • Prediction: Disease Present
  • True Label: Disease Present
  • Result: True Positive

False Positives

A false positive happens when the model predicts positive, but the true label is negative. False positives are incorrect positive predictions.

Disease Detection Example
  • Prediction: Disease Present
  • True Label: No Disease
  • Result: False Positive

Precision Formula

Precision is calculated using true positives and false positives.

Formula

Precision = True Positives / (True Positives + False Positives)

Short Form: Precision = TP / (TP + FP)

The denominator means: All predicted positives.

Precision Calculation Example

Imagine a model makes positive predictions for 100 examples.

Calculation
  • True Positives: 70
  • False Positives: 30
  • Precision = 70 / (70 + 30) = 70 / 100 = 70%

70% of the model's positive predictions were correct.

High Precision

High precision means most positive predictions are correct. It is useful when false positives should be avoided.

Spam Model Example
  • Predicted Spam: 100 emails
  • Actually Spam: 95 emails
  • Normal Emails Incorrectly Marked Spam: 5 emails
  • Precision: 95%

Low Precision

Low precision means many positive predictions are wrong. It can create frustration, risk, or wasted effort.

Spam Model Example
  • Predicted Spam: 100 emails
  • Actually Spam: 40 emails
  • Normal Emails Incorrectly Marked Spam: 60 emails
  • Precision: 40%

Precision in Spam Detection

In spam detection, precision measures how many emails predicted as spam were truly spam.

Example
  • Predicted Spam: 200 emails
  • Actually Spam: 180 emails
  • Precision: 90%

Most emails sent to spam were truly spam — a strong precision score.

Precision in Medical Testing

Precision can be important in medical testing when false positives create stress or unnecessary follow-up.

Example
  • Model Predicts: 100 patients have a condition
  • Actually Have Condition: 70 patients
  • Do Not Have Condition: 30 patients
  • Precision: 70%

30% of positive predictions were false alarms. Recall is also important in medicine because missing actual cases can be dangerous.

Precision in Fraud Detection

Precision matters in fraud detection because false positives can inconvenience users.

Example
  • Fraud Alerts: 1,000 transactions
  • Actually Fraud: 700 transactions
  • Normal Transactions: 300 transactions
  • Precision: 70%

Lower precision means many normal users may have transactions blocked.

Precision in Search Results

Precision is useful in search and recommendation systems because users want relevant results.

Example
  • Search Results Returned: 10
  • Relevant Results: 8
  • Precision: 80%

Most returned results are useful.

Precision vs Accuracy

Accuracy measures overall correctness. Precision focuses only on positive predictions.

Accuracy Asks

Out of all predictions, how many were correct?

Precision Asks

Out of all predicted positives, how many were actually positive?

Precision vs Recall

Precision and recall measure different things. Precision focuses on avoiding false positives. Recall focuses on finding actual positives.

Precision

When the model predicted positive, how often was it correct?

Spam: How many marked spam emails were truly spam?

Recall

Out of all actual positives, how many did the model find?

Spam: How many actual spam emails were caught?

Precision and Recall Tradeoff

Precision and recall often have a tradeoff. If a model becomes stricter about predicting positive, precision may increase, but recall may decrease.

Strict Spam Model

Only marks emails as spam when very confident.

Result: High precision, lower recall. Some spam may be missed.

Less Strict Spam Model

Marks more emails as spam.

Result: Higher recall, lower precision. More normal emails may be marked spam.

When Precision Is More Important

Precision is more important when false positives are costly.

Precision Is Important For
  • Spam detection — avoid sending important emails to spam
  • Hiring screening — avoid incorrectly rejecting qualified candidates
  • Fraud alerts — avoid blocking normal transactions
  • Medical follow-up alerts — avoid unnecessary stress or procedures
  • Search results — avoid showing irrelevant results

Main goal: Positive predictions should be trustworthy.

When Recall May Be More Important

Recall may be more important when false negatives are more costly than false positives.

Recall May Matter More For
  • Disease detection — missing a disease can be dangerous
  • Fraud detection — missing fraud can cause financial loss
  • Safety monitoring — missing a dangerous event can be harmful

Main goal: Find as many actual positives as possible.

Precision in a Confusion Matrix

A confusion matrix shows true positives, false positives, true negatives, and false negatives. Precision uses only true positives and false positives.

Precision Uses
  • True Positives
  • False Positives

Precision does not directly use True Negatives or False Negatives.

Formula: Precision = TP / (TP + FP)

Confusion Matrix Example

This example shows how precision is calculated from a confusion matrix.

Disease Detection
  • True Positives: 45
  • False Positives: 15
  • True Negatives: 30
  • False Negatives: 10
  • Precision = 45 / (45 + 15) = 45 / 60 = 75%

75% of the people predicted to have the disease actually had it.

Perfect Precision

Perfect precision means every positive prediction is correct. However, it does not always mean the model is perfect — it may still miss many actual positives.

Example
  • True Positives: 50
  • False Positives: 0
  • Precision = 50 / (50 + 0) = 100%

The model made no false positive errors, but may have missed many actual positives.

Precision Can Be Misleading Alone

Precision alone does not tell the full story.

Example

A model predicts only 1 email as spam. That email is truly spam.

Precision: 100%

Problem: If there were 999 other spam emails the model missed, the model is not useful.

This is why precision should often be considered with recall.

Precision and F1 Score

F1 score combines precision and recall into one balanced metric. It is useful when both false positives and false negatives matter.

Fraud Detection
  • Precision — How many flagged transactions were truly fraud?
  • Recall — How many fraud transactions were found?
  • F1 Score — Balances both precision and recall.

Precision in Imbalanced Datasets

Precision is often useful in imbalanced datasets because it tells us whether positive predictions are trustworthy.

Fraud Is Rare

Most transactions are normal. Accuracy may be high without finding fraud.

Precision helps answer: When the model flags fraud, how often is it correct?

Improving Precision

Improving precision usually means reducing false positives.

Ways to Improve Precision
  • Make the model more selective
  • Increase the classification threshold
  • Improve feature quality
  • Remove noisy data
  • Fix incorrect labels
  • Reduce false positives
  • Use better training examples
  • Tune the model carefully

Classification Threshold

Many classification models output probabilities. A threshold decides when the model predicts positive.

Example
  • Fraud Probability: 0.82
  • Default Threshold: 0.50
  • If probability is above 0.50: Predict fraud
  • To improve precision: Raise threshold to 0.80

A higher threshold makes the model predict positive only when it is more confident, reducing false positives.

Threshold Tradeoff

Raising the threshold can improve precision, but it may reduce recall.

Higher Threshold
  • Fewer positive predictions
  • Fewer false positives
  • More missed positives
Lower Threshold
  • More positive predictions
  • More actual positives found
  • More false positives

The best threshold depends on the application.

Real-World Example: Email Spam Filter

A spam filter should have high precision because users do not want important emails incorrectly placed in spam.

Spam Filter
  • False Positive — Important email marked as spam
  • High Precision — Predicted spam emails are usually truly spam
  • Possible Tradeoff — Some spam may still reach the inbox

Real-World Example: Job Candidate Screening

A model may predict which candidates should move forward. Precision measures how many recommended candidates are actually strong matches.

Candidate Screening
  • False Positive — A weak candidate is recommended
  • False Negative — A strong candidate is missed
  • Precision Focus — How trustworthy are the recommended candidates?

Recall and fairness should also be checked.

Real-World Example: Product Recommendations

A recommendation system suggests products to users. Precision measures how many recommended products are actually relevant.

Example
  • Recommended Products: 10
  • Relevant Products: 7
  • Precision: 70%

Most recommendations are useful.

Common Mistakes

Common Mistakes
  • Confusing precision with accuracy
  • Confusing precision with recall
  • Looking at precision alone
  • Ignoring false negatives
  • Ignoring the classification threshold
  • Ignoring the real-world cost of mistakes
  • Assuming high precision means the model is always good
  • Ignoring performance across different groups

Summary

Key Takeaways
  • Precision measures how many predicted positives were actually positive.
  • Precision focuses on positive predictions.
  • Precision uses true positives and false positives.
  • Precision formula is TP / (TP + FP).
  • High precision means fewer false positives.
  • Precision is useful when false positives are costly.
  • Precision is different from accuracy and recall.
  • Precision alone can be misleading.
  • Precision is often used with recall and F1 score.
  • Changing the classification threshold can affect precision.

Practice Prompt

A spam detection model marks 80 emails as spam. Out of those 80 emails, 60 are actually spam and 20 are normal emails.

Calculate the precision. Then explain what the precision score means and why false positives matter in this example.

Need Help?

Ask the AI if you need help understanding precision, predicted positives, true positives, false positives, confusion matrices, thresholds, or the difference between precision and recall.