Key Takeaways
- The Pangram AI detector is a general-purpose AI detection system built by Pangram Labs, used across education, publishing, and content teams.
- Pangram publishes strong benchmark numbers. However, in our own testing, Pangram correctly identified fully AI and fully human writing, but struggled with majority-human, ESL, and paraphrased submissions.
- Pangram provides a segment-level output that does not break down results sentence-by-sentence, making it hard to verify what actually drove a given score.
- Documented cases, including inconsistent results a Wall Street Journal columnist got when retesting Pangram’s own study, and user reports of human-written text flagged as AI show that an AI score alone shouldn’t be treated as conclusive proof of authorship.
- In a recent case, Pangram scored Mia Ballard’s Shy Girl novel 78% AI-generated, after which Hachette pulled the UK edition and canceled the US release. Ballard denies using AI herself and is pursuing legal action.
- Proofademic was built specifically for academic review, with a sentence-level score and a paraphrase shield, and performed more consistently across the same tests.
Pangram is an AI detection system developed by Pangram Labs to classify text as human-written, AI-generated, or AI-assisted. The company trains its model on millions of documents, and its published performance claims have also been examined in independent research. However, for teachers, professors, and administrators, the practical question is narrower than the company’s broader claims because these tools can play a role in academic grading and integrity decisions. In this Pangram AI detector review, we examined Pangram’s stated technology and benchmark claims, tested it against academic-style writing, and reviewed independent research and reported user experiences, including cases where human-written work was flagged. We also compared Pangram directly with Proofademic on the same tests, to help academic readers judge which tool actually fits a grading workflow.
Short answer: Pangram scores well on controlled accuracy benchmarks, and has real independent research behind those numbers. However, our testing and reported real-world experiences show that it can still misjudge human academic writing, particularly ESL prose and heavily human-edited drafts. So for grading decisions, treat any single score as a starting point for review.
What is the Pangram AI detector?
Pangram is a machine-learning classification system that scores a document for the likelihood that it was written by a large language model. Pangram markets itself to educators, publishers, recruiters, content creator teams, and enterprises, and it is positioned across all those use cases; it cannot be considered a dedicated academic-integrity tool. Its current release, Pangram 4 (July 2026), adds mixed-authorship detection, an AI-assisted category, and a humanizer-detection signal on top of earlier versions. However, to evaluate its accuracy on real academic submissions, we used the same model to run our tests as mentioned below.
Pangram trains its human-text corpus on licensed long-form writing spanning essays, scientific papers, web texts, books, and professional documents, and generates its AI-text corpus in-house. According to Pangram’s own model documentation, the current model supports detection across languages and is evaluated against generator models from Anthropic, OpenAI, Google, Meta, DeepSeek, and others.
A Pangram score reflects how closely a passage’s statistical fingerprint matches AI-associated patterns. The score is a classification signal and should not be treated as proof of authorship, which also emphasizes why understanding false-positives in AI detection matters before treating any score as conclusive.
What Pangram can do
- Detect fully AI-generated content or text assisted with tools.
- Identify AI involvement at the segment level.
- Detect content from a wide range of AI model families.
- Analyze documents with mixed human-AI authorship.
- Support academic use cases through LMS integration.
What Pangram can’t do
- Cannot provide sentence-by-sentence resolution, as Pangram provides multi-sentence segment-level interpretability.
- It cannot perform equally on every text type, as its documentation identifies certain text types as higher-error domains. Pangram’s own documentation flags short replies, lists, instructions, technical manuals, reference sections, templated writing, and math-heavy text as higher-error domains.
- It is not designed exclusively around academic-integrity workflows.
- The tool cannot prove who wrote the document. The score is just a statistical estimate.
- It cannot replace human judgment in a high-stakes academic decision.
How accurate is Pangram’s AI detector?
Vendor-reported accuracy numbers are useful, but they are measured under controlled conditions. To help students and educators better understand how the tool performs under real academic submissions, we ran a set of academic writing samples through Pangram ourselves, so that our review does not remain confined to benchmark claims alone.
Our methodology
We built five test cases designed to mirror situations that come up constantly in academic settings: a purely AI-generated essay, a fully human-written essay, an essay that was largely human-written, an essay from a non-native English (ESL) writer, and an AI-generated essay that had been paraphrased through a paraphrasing tool. We submitted each sample to Pangram exactly as written and recorded the document-level score Pangram returned for each one. To keep the comparison fair, we tested both tools using the same writing samples. The goal wasn’t to produce a large statistical example, but to check how the tool handles the specific edge cases that actually decide academic-integrity outcomes.
Test 1: Fully AI-generated essay

Result: Correctly identified as 100% AI-generated.
Test 2: Fully human-written essay

Result: For a completely human-written essay, Pangram correctly classified it as 100% human-written.
Test 3: Majorly human-written essay

Result: In our test, Pangram assigned a 100% AI-generated score to a predominantly human-written essay with minor Grammarly and tone adjustments. It raises concerns about the reliability of AI scores when evaluated without supporting evidence.
Test 4: ESL writer’s essay

Result: Pangram returned 62% AI on a non-native English writer submission. This suggests the model struggled to reliably distinguish ESL writing patterns from AI-generated patterns in this sample.
Test 5: Paraphrased AI-generated essay

Result: An AI-generated essay run through a paraphraser returned a result showing it as 57% human (43% AI). In our test, the paraphrasing step was enough to push a fully AI-written essay into a majority-human classification on Pangram.
Testing Proofademic on the same essays
To see how an academia-specific detector handles the same scenarios, we took the same writing samples and ran them through Proofademic. Our methodology remained the same; we used the same essays and the same AI generation and paraphrasing tools to compare both tools.
Test 1: Completely AI-generated essay

Result: The test provided a score of 98% AI, which is accurate.
Test 2: Completely human-written essay

Result: Scored correctly as 97% human.
Test 3: Largely human-written essay

Result: Returned a score of 77% human-written, correctly reflecting that the large majority of the essay was human-authored.
Test 4: ESL writer’s essay

Result: The detector scored the essay as 91% human, correctly identifying human authorship despite non-native phrasing patterns.
Test 5: Paraphrased AI-generated essay

Result: The tool returned a 93% AI score, correctly flagging AI origin even after paraphrasing. This is the type of case Proofademic’s Paraphrase Shield is specifically built to handle, helping identify paraphrased AI-generated content intended to bypass AI detection.
What our tests showed
The table below summarizes how Pangram and Proofademic performed across the same five writing scenarios, using samples with known origins.
| Ground truth | Pangram | Proofademic |
|---|---|---|
| Fully AI-generated | 100% AI | 98% AI |
| Fully human-written | 100% human | 97% human |
| Majority-human written | 100% AI | 33% AI |
| ESL human | 62% AI | 91% human |
| Paraphrased AI | 43% AI | 93% AI |
Across five scenarios, Pangram’s results were less consistent across the more complex scenarios we tested, such as majority-human writing, ESL prose, and paraphrased AI content. We also noticed that three of our five Pangram tests scored at an extreme 100% classification, providing no nuanced and proportionate scores even on content that combined human and AI characteristics. Also, Pangram’s segment-level output does not break results sentence-by-sentence, so we had no way to verify which specific sentences were driving a given score.
On the other hand, across the same five scenarios, Proofademic’s results tracked each sample’s known origin more closely than Pangram’s did, particularly on the majority-human, ESL, and paraphrased samples. Proofademic also provided a sentence-level heat map showing which specific sentences contributed to the overall score, which meant we could see precisely what drove each result.
Can Pangram produce false-positives on human writing?
Pangram’s published figures report a 0.0041% false-positive rate on its Pangram 4 evaluation benchmark, and the University of Chicago’s Becker Friedman Institute found that Pangram met the false-positive policy threshold in controlled testing. However, academic writing does not usually follow a controlled pattern, and the online reviews and independent testing tell a different picture.
Shy Girl By Mia Ballard Case
The Shy Girl case is the most prominent example that shows how an AI detection score can have consequences beyond content moderation. Suspicion around the novel started with a Reddit post and a YouTube video essay questioning the prose, before Pangram entered the picture at all. Pangram’s detector later returned a 78% AI-generated score on the book, which Spero shared publicly. The score became part of a wider news story that reached Hachette Book Group, which pulled the UK edition and canceled the US release following its own investigation. Ballard has denied using AI and said the incident has damaged her reputation.
Score Discrepancy Over Time
Independent scrutiny has also raised questions about the consistency of Pangram’s results. Writer Tim Requarth documented an episode in which Wall Street Journal columnist James Taranto independently retested Pangram’s AI detection results by re-running the three op-eds previously with AI-generated flags in a Pangram study through the company’s own public tool. The results reportedly differed from the earlier classifications, as when he re-ran the results were inconsistent, with one article reportedly receiving a 100% human score despite having previously been identified as AI-generated. Requarth described the results as serious and cautioned against treating a single score as conclusive.
Reviews from Real Users
Alongside these benchmark results, anecdotal reports from users and current Reddit reviews of Pangram provide a different perspective on how the detector can behave in individual real-world cases. In a Reddit discussion among academics and writers, several users described human-written work being flagged, including one writing professor who said Pangram identified his work as “AI-generated, with a 100% confidence,” while another user complained that “I was flagged for 100% AI on Pangram, but the report also said low confidence, which doesn’t make sense to me.” The same user further stated that they removed bullet points and headings; the exact same paper came back as 100% human-written.
While these individual experiences do not determine Pangram’s overall false-positive behavior by themselves, they do offer useful insight into how AI detection performs in real-world academic and professional writing.
What should educators do with a Pangram flag?
A detection score should initiate a conversation or any review process. Before treating any AI-detector result as evidence, a stronger review process holds up better under scrutiny, and a workflow like the following gives an educator a defensible basis for a decision.
- Review the flagged passages directly before relying on the headline percentage alone.
- Compare the submission with the student’s previous work to check for consistency in voice, vocabulary, and structure.
- Check drafts and version history, where available, since a genuine writing process usually leaves a trail.
- Review citations and sources for consistency with the flagged sections.
- Consider the assignment and course context, including if AI-assisted tools like grammar checkers for editing were permitted under course policy.
An AI detector can flag patterns for review, while the full context helps educators reach a fair and informed conclusion. This principle is central to responsible AI detection in academic settings and is one reason institutions are increasingly comparing AI detectors colleges use in 2026 when developing broader AI and academic-integrity policies.
Is Pangram free? Pricing and practical value
Pangram offers a free tier with a daily usage limit and a capped word count, according to its current pricing page. The limited word count available in the free tier is workable for spot-checking a paragraph, but if you want to check a full essay, dissertation chapter, or thesis, it exceeds what a free plan is designed to handle in a single check. An important limitation to consider here is that Pangram does not provide access to its plagiarism checker without a paid plan. So if you want to run an academic submission for AI and plagiarism checks in a single scan, it is not possible on a free plan; paid individual and professional plans raise the monthly ceiling.
Pangram vs. Proofademic: Which is the better fit for academia
Proofademic and Pangram are both capable AI detectors, but they were built for different jobs, and that distinction is what should actually drive a decision for academic use. Based on the vendor documentation, published research, and our own testing above, here’s how the two compare on the criteria that matter most in a classroom or institutional setting.
| Criteria | Pangram | Proofademic |
|---|---|---|
| Positioning | Broad detection platform serving education, publishing, recruiting, and content generation teams together | Purpose-built specifically for students, instructors, and publishers in academic settings |
| Result interpretation | While testing, we saw that its current model produces segment-level output, in place of sentence-by-sentence detail | Sentence-by-sentence scoring with a visual heatmap showing exactly which lines drove the result |
| Accuracy evidence | Strong published and independently validated benchmark data, but in our own testing results were less consistent on edited, ESL, and paraphrased submissions | Strong results across our test set, including on edited, ESL, and paraphrased AI content, aided by a dedicated paraphrase shield built to catch reworded AI text |
| Academic fairness | Documentation acknowledges a non-zero error rate and takes false-positive reports seriously, but the product itself is not built exclusively around academic review workflows | Fairness, sentence-level transparency, and integrity review are the core of AI detector, alongside a combined plagiarism and AI check |
| Best fit | Organizations that want Pangram’s established detection ecosystem, multi-industry integration, or broad enterprise use | Students, instructors, and publishers who want an academia-first detection experience with interpretable, sentence-level review |
The comparison between Proofademic and Pangram shows that AI detection for academia is a narrower and more specific problem than general-purpose AI detection, and Proofademic’s academic AI detector is built around that narrower problem directly: interpretability, fairness, and a workflow that matches how grading decisions actually get made.
Why is Proofademic a better tool for academic use cases?
Grading decisions carry real consequences for students, which is exactly why a detector built specifically for classrooms matters more than a generic accuracy claim. Proofademic was designed from the ground up for academic AI detection: a plagiarism checker and AI text detector running in one workflow, sentence-level scoring to help instructors precisely see what drove a result instead of an unexplained percentage, a paraphrase shield built specifically to catch reworded or paraphrased AI text, multilingual detection, and batch scanning for departments reviewing large volumes of submissions.
In our side-by-side testing, this showed up directly in the results as Proofademic correctly scored the majority human essay, the ESL writer’s submission, and the paraphrased AI essay; the three scenarios where Pangram diverged most from the actual writing. Educators who prioritize explainability and sentence-level transparency may find its workflow better suited to academic integrity decisions. Try it with a 3-day free trial, no credit card required, before deciding if it fits your workflow.
TL;DR
Pangram is a capable AI detector with real independent research behind it. But in our own testing, we found it struggled with the situations that matter most in a classroom, such as majority-human writing, ESL prose, and paraphrased AI content. Also, reported real-world false-positive cases show why a single score should not be the final word on an academic-integrity decision. The right decision for educators is to treat any detector output as a starting point for review. In our five-sample comparison, Proofademic matched the known origin of the majority-human, ESL, and paraphrased samples more closely with sentence-level interpretation, while Pangram has the publicly available independent large-scale benchmark record.
FAQs
Is Pangram a good AI detector?
Pangram has real third-party validation behind its published numbers, making it a credible detector, but good aggregate performance does not make any single classification definitive proof of authorship for an individual document.
Is Pangram’s AI detector free?
Pangram does offer a free plan with limited word count and usage limits. However, it is a better fit for short texts and is generally insufficient for a full academic submission.
How accurate is Pangram’s AI detector?
Pangram’s own published benchmarks and independent academic research report a very low false-positive rate under controlled test conditions. However, that figure is distinct from how the tool performs on individual real-world submissions, which can vary meaningfully depending on the writing scenario. Our own testing identified several situations where Pangram’s classification did not align with the known origin of the sample.
How does Pangram AI detector work?
Pangram uses a trained classification model that compares statistical patterns in submitted text against patterns learned from human and AI-written training examples, then returns a score and labeled segments.
Can Pangram’s AI detector produce false positives?
Yes. AI detectors, including Pangram, are fallible, and Pangram’s own documentation acknowledges a non-zero error rate. Our testing and current user reports both describe cases of human academic writing being misclassified, though these should be understood as anecdotal accounts rather than a formally measured false-positive rate.
How to bypass Pangram AI detector?
No ethical approach should focus on bypassing or manipulating Pangram’s AI detector to disguise AI-generated work as human-written. Instead, students and institutions should follow applicable academic-integrity policies and be transparent about how AI tools were used. If a human-written submission is incorrectly flagged, supporting evidence such as draft history, notes, version records, and other documentation can help establish authentic authorship.
How does Pangram compare to Proofademic?
Pangram has an established detection footprint across industries and a strong published record. Proofademic is built specifically for academic detection, with sentence-level interpretation, a combined plagiarism and AI workflow, and fairness and integrity as core product goals rather than add-ons.


