DaidaRedact Help
Run a redaction batch, understand where your records go, and recover work after an interruption.
The current app uses your browser to coordinate batches and combine large PDFs. Closing the tab, a sleeping device, or an expired session can interrupt that work. Azure jobs already accepted continue independently. Opening Help from the app uses a separate tab.
1. Start a batch
- Connect Azure. Use a single-service Azure AI Language resource and separate source and output Blob containers. Enter the Language endpoint/key and container SAS access links. The app checks the key with a small text request and checks storage access; this does not prove that document black-out processing will succeed.
- Add records. Select local PDFs or import from an AWS S3 bucket/folder or Azure Blob location. A file may lack a
.pdfextension, but its contents must be a readable PDF. This app submits documents as English. - Check pages and cost. Wait for preparation to finish. The estimate uses the page count and your entered rate per 1,000 pages. For a planning rate of $0.01 per page image, enter $10 per 1,000 pages; a 10-page PDF estimates at $0.10. This is your configured estimate, not an Azure billing quote; storage, hosting and repeat processing can add cost.
- Upload / import, then choose entities and style. New batches start with no entities selected. For names and phone numbers, choose Person and PhoneNumber. Black out uses Azure’s preview document PII policy. Other choices are asterisks, entity labels and synthetic replacement; the app never silently substitutes a different style.
- Process and review. Review the estimate, then choose Process batch. Saved settings apply throughout that batch. After processing, review the redacted outputs and prepare the comprehensive batch report.
2. Sessions, interruptions & recovery
It can. Combining downloads redacted parts, merges them in your browser, then uploads and saves the combined PDF. Each transfer requires an active session. A calculation already in progress may finish, but the next protected request can fail and send you to sign-in. An open tab alone is not enough.
What continues or survives
Azure jobs already accepted continue. Fully uploaded PDFs, saved batch progress, redacted parts and completed combined files remain in Azure until deleted or expired under your storage rules.
Files only queued in the browser are not yet saved. An interrupted upload may need the original file selected again. An interrupted combine may restart for that PDF using its saved redacted parts.
How to resume
- Sign in again.
- Reconnect to the same Language resource and source/output containers. Renew expired SAS links if needed.
- Choose Open batch and select the saved batch. Your last batch may reopen automatically.
- Choose Resume batch. Completed redacted parts are reused for combining.
Use Retry rejected groups only for submissions Azure definitively rejected. A submission with an unknown outcome may already be running; inspect it before retrying.
Hosted edition
Sign-in uses ChatGPT and the hosting service. This app does not define that service’s session timeout. Expired or revoked access stops subsequent protected requests.
Windows / IIS edition
The packaged app uses email/password accounts. Sessions expire after 30 minutes without authenticated requests and after eight hours total. Normal processing requests refresh the idle timer; mouse inactivity by itself does not end the session. The eight-hour limit still applies.
Neither edition currently includes an independent background batch worker. The Windows service hosts the web app; it does not remove the browser dependency. Keep this distinction in mind when planning overnight or unattended jobs.
3. How records move through the app
Follow one original record from import to delivery. Large records may create several processing parts, while the batch report keeps their relationship to the original.
- Your device or cloud sourceSelect the original PDF
Choose local files, S3 objects or Azure blobs. Cloud originals are read and copied, not moved.
- Browser / app serverValidate, count and estimate
Local PDFs and large cloud PDFs are prepared in the browser. Smaller cloud PDFs are inspected by the app server. Cost counts original pages once.
- Browser → app server → source storageSave originals and input parts
PDFs over 10 MB are split on page boundaries. The original is kept intact. Each part must fit the service limit.
- Browser → app server → Azure LanguageSubmit groups with your settings
A group contains at most 40 documents and 10 MB combined. New groups can submit while accepted jobs run; Azure throttling triggers a wait.
- Azure Language → output storageDetect and redact
The service reads the inputs and saves redacted files and detection JSON. Accepted jobs continue even if the browser closes.
- Output storage → app server → browserCombine large records
After every required part succeeds, the browser downloads the redacted parts through the app and combines them in original page order.
- Browser → app server → output storageSave the final redacted PDF
The app verifies the complete transfer before making the combined output available. Missing or failed parts prevent a final combined PDF.
- App / source storage → reviewerReport, review and deliver
Prepare a batch summary and PSV or JSON detection export. Review outputs before sharing; then apply your retention policy to originals, parts, outputs and reports.
4. Architecture & technology stack
The two deployment editions use the same redaction workflow. Their web hosting and user-account storage differ.
Choose entities • review cost
Prepare / split PDFs • drive batch progress
Combine redacted parts • prepare reports
Check account access on each request
Read cloud sources • transfer files
Submit jobs • check status • save progress
Original PDFs, smaller input parts, batch records and batch reports
Document-based PII detects and redacts selected entities in accepted jobs
Redacted PDFs / parts, raw detection JSON and final combined PDFs
| Layer | Hosted edition | Windows / IIS package |
|---|---|---|
| User interface | React + TypeScript. Browser Web Worker and pdf-lib prepare, split and combine PDFs. | |
| App server | Next.js-style app routes built with Vinext/Vite, running in a Cloudflare Worker through Sites. | Next.js on Node.js as a Windows service, behind IIS HTTPS reverse proxy. |
| Login & user administration | ChatGPT sign-in; application roles and approval state in Cloudflare D1. | App-managed email/password accounts, roles and sessions in SQLite on the server. |
| Redaction service | Azure AI Language, Document-based PII, through its asynchronous REST API. The app does not make a separate Azure Document Intelligence call. | |
| Document & batch storage | Your Azure Blob source and output containers. Optional AWS S3 and Azure Blob import sources are read through the app server. | |
| Batch coordination | The browser requests work; the server uses saved Blob records and leases to coordinate submissions. There is no independently running queue worker. | |
The app currently uses API 2026-05-01 for standard policies and 2026-05-15-preview for black-out and synthetic replacement. Every new submission and retry requests model version “latest”, while retaining the batch’s API version, style and entities. This is a compatibility adjustment; a small pilot is still needed to confirm black-out processing on your resource.
5. Storage, reports & retention
Processing source container
Original PDFs, split input parts, batch/job records, report preparation data and combined PSV/JSON report exports.
Output container
Azure’s redacted files and detection JSON, plus final combined PDFs saved by the app. Both containers are in your Azure account.
Browser and accounts
Azure connection details stay in the tab’s memory and are sent to the app server for requests. Reopening or signing in again can require reconnecting. The browser also temporarily holds selected files and PDF parts.
Account databases hold application access information. They are separate from the Azure containers that hold your documents.
PSV and JSON exports may include the detected names, phone numbers, categories and confidence scores. Treat them as sensitive records, even when the output PDFs are redacted. Reports record the requested model selector for each processing part. Older jobs without that record are marked unknown. The original saved model preference is listed separately; “latest” is a selector whose resolved Azure model may change over time. Prepare a fresh report after upgrading to include these fields. A report’s low-confidence threshold helps prioritize review; it does not change what Azure redacted. No detections is not proof that a document contains no PII.
The app does not automatically delete your Azure files. Configure storage lifecycle rules or delete records under your retention policy. Keep required parts and batch records until processing and review are complete. Azure processing, storage, transactions and any transfer charges are billed to your subscription.
6. Troubleshooting & practical limits
Azure rejected the request with HTTP 400
Use Copy error details below the failed file. The message includes a filtered Azure explanation when supplied, the rejected field when available, the request ID and the selected processing settings. Do not paste keys or SAS links. Use one small test file to investigate before processing a larger batch. A resource connection test alone does not validate black-out redaction.
Azure is busy, or the batch appears paused
The app respects Azure’s retry delay. Keep the tab open and signed in. If the app paused because of a session or network problem, reconnect and resume. It does not automatically replay a submission whose acceptance is uncertain.
A single page is larger than 10 MB
Splitting at page boundaries cannot make that page smaller. Reduce the scanned image size in an approved PDF tool, verify legibility, then try again. The app does not silently reduce image quality.
How many files can I process?
The current app accepts up to 5,000 original PDFs per batch. Each Azure request is grouped to 40 documents or 10 MB combined, whichever is reached first. The app limits work within each HTTP step but does not cap the number of Azure jobs already accepted. Browser memory, network speed and Azure capacity affect turnaround; add smaller batches if your device struggles.
How do I know which version I am using?
Look for v1.2.2 beside Help in the app navigation, on the sign-in page, or at the top of this guide. Refresh after an update. When reporting a problem, include the version and Copy error details text.
Service reference: Microsoft’s document PII guide. Published service limits and feature availability can change; this help describes DaidaRedact 1.2.2.