Automated Data Classification for Every File
Label PII, PHI, and other content across SharePoint, email, file shares, and more.
Shinydocs classifies files according to sensitivity and value to improve findability, access control management, and retention schedule enforcement across systems.
- Of privacy leaders are confident their organization can ensure the privacy of its sensitive data.
- 43%
Source: ISACA, State of Privacy 2026
- Of organizations lack the right data management practices for AI, or are unsure they have them.
- 63%
Source: Gartner, 2025
Results
Proven Data Classification Results
Shinydocs customers classify millions of files for organizations to help them reach 100% compliance across four million files and save teams time, without migrating any content.
- FOI compliance
- 100%
- Hours saved weekly
- 6,500
- Documents governed
- 98%
Town of Milton reached full FOI compliance across four million files. Source: Town of Milton case study
Dunedin City Council reclaimed 6,500 staff hours every week. Source: Dunedin City Council case study
County of Newell brought 98% of its documents under governance. Source: County of Newell case study
Data classification explained
What Is Data Classification?
Data classification organizes files into categories based on sensitivity, value, and business importance. It applies to both structured and unstructured data. Common levels include public, internal, confidential, and restricted. Organizations classify manually, automatically, or through a combination of both. Classificaiton helps manage access controls, retention schedules, and defensible disposal across business systems.
Shinydocs' automated information governance software classifies files where they live, with no migration required.
-
π·οΈ
Sensitivity Levels
Sorts content into tiers. Public, internal, confidential, and restricted each carry different handling rules.
-
π
Structured and Unstructured Data
Covers database records alongside the documents, emails, spreadsheets, and other files teams create.
-
π
PII and PHI Detection
Finds personal identifiers, health records, financial data, and credentials inside ordinary files.
-
π
Labels and Metadata
Every classified file carries a label. Security tools, retention schedules, and AI systems all read that label.
-
π
Content and Context Signals
Classification reads what a file contains and where it came from. Both signals decide the label.
-
π
Access, Retention, and Disposal
Classification labels determine who can open a file, how long to keep it, and when it gets deleted.
Discovery to review
How Data Classification Works
Data classification runs continuously rather than as a one-time project. Newly created and updated files are classified on an ongoing basis. You define the objectives and sensitivity levels once. Shinydocs discovers content across every connected repository, applies labels, and enforces the controls each label carries. Scheduled reviews confirm the scheme still fits your regulations and business use.
- 01
Define Objectives
Decide what needs protection and why. Regulations, contracts, and risk appetite set your categories.
- 02
Discover Content
Scan every repository in place. Find what exists across file shares, email, and cloud drives.
- 03
Set Sensitivity Levels
Define each tier in plain language. Clear definitions keep labels consistent across every team.
- 04
Apply Labels
AI classifies at volume. Reviewers decide only ambiguous files, and every action is logged.
- 05
Enforce and Review
Controls follow the label. Scheduled reviews catch drift as content and regulations change.
What the software does
Data Classification Capabilities
Our AI data classification platform finds sensitive content, labels it, and keeps those labels current. Shinydocs reads what each file contains and where it came from, then writes the label back into metadata. Existing permissions hold through the process, and every decision lands in an audit trail.
| Capability | What It Does |
|---|---|
| Automated Discovery | Scans every connected repository and inventories what content exists. |
| Content Analysis | Reads inside files to detect patterns, identifiers, and sensitive terms. |
| Context Analysis | Infers sensitivity from source system, owner, location, and file age. |
| Sensitive Data Detection | Flags personal, health, financial, and credential data wherever it sits. |
| Metadata Enrichment | Writes labels and attributes back so other systems can act on them. |
| Permission-Aware Processing | Respects existing access controls throughout classification and review. |
| Continuous Reclassification | Re-evaluates content as it changes, so labels stay accurate over time. |
| Audit-Ready Reporting | Records every classification decision for regulators and internal auditors. |
Sensitivity tiers
Data Classification Levels
Organizations typically classify data across four sensitivity tiers: Public, Internal, Confidential, and Restricted. These tiers determine which controls apply based on the level of risk. Archived content remains subject to applicable retention requirements and must be disposed of defensibly when retention expires.
| Level | What It Covers |
|---|---|
| Public | Press releases, marketing material, and published research. Disclosure causes no harm. |
| Internal | Employee directories, internal memos, and training material. Exposure causes inconvenience, not legal liability. |
| Confidential | Customer lists, contracts, financial reports, and pricing. Disclosure causes financial or reputational harm. |
| Restricted | Health records, payment card data, account numbers, biometric identifiers, and trade secrets. Disclosure creates legal exposure. |
Four approaches
Types of Data Classification
Organizations classify data four ways. Content-based classification, cntext-based classification, user-based classification, and automated classification. Mature programs typically combine several methods rather than depending on just one. We outline all four below. Shinydocs incorporates each one.
-
π
Content-Based Classification
Scanners read the file and match patterns. Identifiers like health record numbers and payment card data trigger a label.
-
πΊοΈ
Context-Based Classification
Rules infer sensitivity from metadata. A file from a payroll system inherits a confidential label automatically.
-
π€
User-Based Classification
People tag their own files. Human judgment catches nuance, but the method rarely scales past a few teams.
-
π€
Automated Classification
Machine learning labels content at volume. High-confidence results apply directly, and ambiguous files route to a reviewer.
The comparison
Manual vs. Automated Data Classification
Manual classification requires employees to tag their own files. Accuracy depends on attention and consistency, and it degrades as volume grows. The work is time consuming, and most teams can't spare the hours. Automated classification applies rules and machine learning across every connected system at once. Millions of files receive accurate labels without added headcount.
| Manual Classification | Automated Classification |
|---|---|
| Staff tag files one at a time | Every connected system classified at once |
| Labels drift as content is updated | Reclassification runs continuously for accuracy |
| Coverage stops at employees who comply | Coverage reaches every file in scope |
| Backlogs grow faster than teams can clear them | Millions of files process without added headcount |
| Audit evidence lives in spreadsheets | Every classification action is backed by an audit trail |
Why it matters
Why Does Data Classification Matter?
Data classification supplies the labels that security and compliance rules need. A file marked 'restricted' inherits tighter permissions, a specific retention period, and closer monitoring. Labels drive access reviews, where teams confirm who can open each file. They also answer subject access requests, when someone asks what data you hold. Without them, retention schedules and disposition rules have nothing to act on.
-
Treating every file as equally sensitive wastes budget and slows everyday work. Classification directs strong controls at restricted content and lets low-risk content move freely.
-
You cannot protect data you have never located. Classification produces an inventory of where personal, health, and financial content actually sits.
-
Permissions granted years ago rarely match who needs the file today. Labels give security teams the evidence to tighten access where sensitivity demands it.
-
Most breaches expose content nobody knew was there. Classification shrinks that surface by finding sensitive files early and flagging them for protection or removal.
-
Access requests and subject requests carry statutory deadlines. Classified content lets teams locate responsive records in minutes rather than searching each repository by hand.
-
A retention schedule only works when the system knows what each record is. Classification supplies that answer, so schedules apply automatically from creation forward.
-
Organizations keep content indefinitely because deleting it feels riskier than storing it. Classification supplies the evidence that makes defensible deletion a documented decision instead of a gamble.
-
AI assistants surface whatever they can reach, including content they should not. Classification marks sensitive material before it reaches a model. McKinsey reports that only 7% of companies have fully scaled AI, with data readiness the constraint. Source: McKinsey, 2026
-
Redundant, obsolete, and trivial content sits in paid storage for decades. Classification identifies what carries no business value, so you stop paying to keep it.
-
Auditors ask which protections apply to which content, and how you know. Classification records every decision, so the answer is a report rather than an assurance.
-
Different jurisdictions impose different rules on the same category of record. Classification captures origin and content type together, so the right rule applies to each file.
-
Every department interprets a classification policy slightly differently. Automation applies one scheme everywhere, so a contract carries the same label in legal and in finance.
Where teams use it
Data Classification Use Cases
Privacy teams locate PII, PHI, and payment card data, then answer subject access requests. Legal teams identify client matters, run eDiscovery, apply holds, protect intellectual property, and separate records during divestitures. Records teams assign retention, dispose defensibly, answer FOI and ATIP requests, and prepare for audits. IT scopes migrations, readies content for AI, fixes oversharing, and clears legacy shares.
-
Find personal information sitting unprotected across file shares and cloud drives. Teams reduce PIPEDA exposure before regulators ask.
-
Identify patient data, research records, and regulatory submissions across clinical repositories. A hospital might classify a record as HIPAA-restricted while granting different access to clinicians and administrators.
-
Tag unclassified files to the correct matter number across document management systems and file shares. Misfiled records stop accumulating.
-
Answer subject access and erasure requests without manual repository sweeps. Labels show exactly where one person's data lives.
-
Apply approved retention schedules to every classified record automatically. Your policy becomes enforced reality instead of a documented ideal.
-
Apply disposition authorities from Library and Archives Canada and NARA to classified records. Every action captures an audit trail.
-
Classify content before a platform move. Teams see what deserves to travel and what should be disposed of first.
-
Mark sensitive and obsolete content before AI tools index it. Assistants then draw on governed material rather than everything available.
-
Compare sensitivity labels against who can currently open each file. Security teams see exactly where permissions exceed the content's risk.
-
Produce classification reports on demand for examiners. Organizations demonstrate governance without a last-minute scramble across systems.
-
Identify privileged and responsive material early in a matter. Holds then apply to classified records rather than whole folders.
-
Classify decades of unmanaged shared drives. IT reclaims storage knowing which files carry value and which do not.
-
Locate designs, formulations, and proprietary research scattered outside controlled systems. Restricted labels then bring those files under proper protection.
-
Find cardholder data sitting outside the systems your PCI DSS scope documents. Assessors receive an accurate boundary rather than an assumed one.
-
Separate records by entity, matter, or department during a transaction. A complete, classified file follows the business that owns it.
Extraction, models, and confidence
What Is a Classification Model?
A classification model is the logic that decides which label a file receives. Your classification scheme defines the categories; the model applies them. It reads both what a file contains and where it came from, then scores how well each category fits. That score determines whether the label applies automatically or reaches a reviewer.
-
The input
Text Extraction
Shinydocs pulls readable text from documents, spreadsheets, email, and scanned images. Classification reads content, not file names.
-
The match
Patterns and Models
Fixed patterns catch structured identifiers such as health record and card numbers. Models score the remaining content against your categories.
-
The decision
Confidence and Review
AI Classifiers can be used as validation. Results are specific to the training and test set for prediction of confidence or score.
-
The output
Labels Written in Place
Shinydocs records each file's label without altering the file or moving it from where it lives.
From label to outcome
Classification and Defensible Disposition
Defensible disposition means you can prove the decision was correct. Classification supplies the two facts that matter most: what the file contained, and which rule applied to it. Shinydocs records both at the point of classification, so the evidence exists before anyone deletes anything. See automated data remediation for how disposal runs at scale.
Without classification
With classification
-
β
Nobody knows what the file contained
βEvery file carries a recorded sensitivity label
-
β
Deletion rules apply by folder, not content
βRetention rules apply by what the file holds
-
β
Legal holds depend on somebody remembering
βHolds apply to classified records automatically
-
β
Disposal stalls because the risk is unknown
βDisposal proceeds against an approved schedule
-
β
Auditors receive assurances
βAuditors receive a complete audit trail
Regulatory coverage
How Data Classification Supports Compliance
Every privacy, records, and payment regulation asks the same three questions. What personal data do you hold, where does it live, and who can reach it? Classification helps answer all of them. It identifies regulated content under privacy law, health law, payment standards, and freedom of information law. The correct handling rules then apply automatically, and every decision is documented.
Regulations and What Classification Addresses
| Regulation | What Classification Addresses |
|---|---|
| GDPR | Locate personal data, support erasure requests, and evidence lawful processing. |
| HIPAA | Separate protected health information so it receives the strongest controls. |
| CCPA | Trace consumer data and answer access or deletion requests on time. |
| PIPEDA | Identify personal information across repositories and limit unnecessary exposure. |
| PCI DSS | Scope payment card data so the cardholder environment stays defined. |
| ATIP | Find responsive records across systems within statutory response deadlines. |
| FOIA | Locate responsive records and flag exempt content before public release. |
Legislation reviewed: August 2026
Getting it right
Data Classification Challenges and Solutions
Most classification programs fail for the same reason: they depend on people to maintain them. Labels drift, tags go stale, and new systems appear outside the program. Every one of those failures traces back to the same dependency. Automation removes it, and clear ownership keeps the scheme current as the business changes.
Challenge
Solution
-
Teams label the same content differently
βOne scheme applied everywhere
-
Sensitivity changes but labels do not
βContinuous reclassification
-
Regulations shift faster than internal policy updates
βContinuous alignment with current law
-
Unmanaged systems sit outside the program
βEvery repository connected
-
Manual tagging cannot match file volume
βAutomated labelling at scale
-
Nobody owns the classification scheme
βReporting that assigns accountability
-
Over-classification blocks everyday work
βTiers calibrated to real risk
Works with your stack
Classify Content Across Every Connected System
Shinydocs classifies content where it lives, across existing systems. Connectors index and label each source in place, so one policy covers every repository. Nothing moves, nothing gets copied, and your teams keep working in the tools they know. Coverage spans cloud storage, email, file shares, and document management platforms.
Repository and System Coverage
We add new connectors regularly. Don't see yours? Ask us β coverage keeps expanding.
Data Classification Is One Part of the Platform
Shinydocs is 4-in-1 automated information governance software. It searches, classifies, cleans, and governs content without migrating files. Classification is what the other three run on. Search needs an index, remediation needs to identify ROT, and retention needs to know which schedule applies. One deployment covers all four.
-
π
Enterprise Search
Find any file across every connected repository from one search bar.
Explore β -
π·οΈ
AI Data Classification
Label content with high-confidence AI, so governance and AI tools stay accurate.
You are here
-
π§Ή
Data Remediation
Dispose of redundant, obsolete, and trivial files with a full audit trail.
Explore β -
π
Information Lifecycle Management
Apply retention schedules automatically, from creation through to disposition.
Explore β
FAQ
Frequently Asked Questions
Data Classification in Action
On a quick call, we'll map your repositories and show you what Shinydocs finds inside them.
Book a Meeting
Eliminate Duplicates & Protect Sensitive Data
Automatically detect and filter out redundant files and Personally Identifiable Information (PII), ensuring AI models process only clean, compliant, and relevant data.
AI-Driven Compliance & Governance
Stay compliant by ensuring AI models never process restricted dataβgiving you confidence in security, governance, and regulatory alignment.