Skip to main content
PDF Tools · Security & Privacy

Your documents stay yours

EqualWeb's document remediation runs on a proprietary multi-stage engineering pipeline - semantic structure reconstruction, font-level Unicode repair, and an 80+ rule validation engine - architected so that protecting your data is not a policy layered on top, but a property of the design. Here is exactly what leaves, what is kept, who decides, and what we will never do with your files.

  • No AI training on your files - ever
  • Retention you control
  • ISO/IEC 27001 certified
The one-table answer

Straight answers for your security review

QuestionAnswer
What leaves the system during processing?Only the minimal visual snippets required for the accessibility work - never your document's text layer as a corpus, no metadata, no file names, no client identifiers.
When is a whole document transmitted?In exactly two cases: conversion to HTML, or a scanned file with no text layer that requires OCR - and in both, processing happens in a secured environment with nothing kept beyond the job.
Is anything used to train AI models?No. Never. We do not train models on client documents, do not fine-tune on them, and do not retain documents, content fragments or images for training, model improvement or product development.
How long are documents kept?You decide. Retention is set in your account settings - 0 days, 90 days, 360 days, or never delete - and the system deletes automatically on schedule.
Is data sold or shared?No. There is no onward transfer to additional third parties, no sharing for marketing purposes, and no data monetization.
Is there human oversight?Yes - a client review stage before remediation is performed, and human approval for every piece of AI-generated content, presented next to a screenshot of the exact region.
Data minimization by design

The minimum-exposure principle

At the core sits a proprietary deterministic structure engine that reconstructs the document's entire semantic architecture - tag trees and role mappings, heading hierarchies, table header association, logical reading order, artifact isolation, and font-level Unicode repair - entirely in our own code, on our own infrastructure, with nothing transmitted anywhere. Only the narrow class of operations that require genuine content understanding - such as phrasing a human-quality description for an image - engages content-aware AI vision analysis through hardened commercial interfaces. Each such operation receives a cryptographically isolated visual fragment: the exact region in question, and nothing more. Your document's text layer is never transmitted as a corpus, and no metadata, file names or client identifiers ever travel with a request.

A complete document is transmitted in exactly two scenarios: conversion to HTML, or a scanned file with no text layer that requires character recognition - both in a secured processing environment, with nothing retained after the job completes.

Your data, your schedule

Retention you control

How long processed documents are kept is your decision, set in your account settings - and the system deletes automatically on your schedule. Accessibility reports are retained so your permanent report links keep working, under the same control.

0days - delete immediately after processing
90days - then automatic deletion
360days - then automatic deletion
never delete - for institutions that require continuous records

Model training - an explicit commitment: we do not train AI models on client documents, do not fine-tune models on them, and do not retain documents, content fragments or images for training, model improvement or product development. Access to AI engines is through commercial interfaces under contractual terms.

People, not just pipelines

A human in the loop, by design

Review before remediation

Fast preliminary audit → your review and approval of the scope → remediation → re-validation with visual evidence. A person approves the work before any document is altered.

AI content is a recommendation

Image descriptions, table summaries, field labels and link descriptions are presented in the report next to a screenshot of the exact region - for fast human approval, especially in documents of record.

An audit layer that hides nothing

Every remediated document is re-audited by our validation engine - 80+ rules derived directly from the ISO and WCAG specifications, spanning structure, semantics, encoding and contrast. A finding that cannot be fixed automatically is reported transparently - never concealed behind a value that merely looks conformant.

Infrastructure security

Encrypted, certified, permission-gated

  • End-to-end transport encryption (HTTPS/TLS) at every stage of the pipeline
  • Keys and credentials stored encrypted (AES-256) - never exposed to a browser, never logged
  • Per-action account permissions: who checks, who remediates, who manages settings
  • No onward transfer to additional third parties, no marketing use, no data monetization
ISO/IEC 27001GDPRCCPA
Measured against the standards that matter

Every document, validated against global standards

PDF/UA-1 (ISO 14289-1)PDF/UA-2 (ISO 14289-2:2024)WCAG 2.1WCAG 2.2ADASection 508EN 301 549

EN 301 549 explicitly covers non-web documents - your PDFs are in scope in their own right, wherever in the world you operate.

Security review?

Bring your toughest questions

We answer vendor assessments, DPAs and security questionnaires as part of onboarding - and we would rather show you how it works than ask you to take our word for it.

Talk to a specialistBook a meeting