NearScrub

2026-08-16

What your PDF or Word attachment reveals when you file a public comment with a federal agency

Filing an official comment on a proposed federal rule often means uploading a file rather than just typing into a web form — a formal letter, a technical comment, a spreadsheet of supporting data. That file gets published to the government's public docket exactly as it was uploaded, and several federal agencies say so in their own guidance. What none of that guidance mentions is what else typically travels inside a Word or PDF file besides the text on the page.

What agencies actually promise about your attachment

The U.S. Fish and Wildlife Service's page on how it handles public comments is explicit about the file itself, not just the comment box: "If you submit a comment, your entire comment—including any personal identifying information—may be available to the public," and FWS offers "no guarantee" of withholding personal information "if you include any personally identifiable information in the comment text, as part of an attachment (MS Word doc, Adobe PDF, MS Excel, etc.)." The CFTC's instructions for submitting a comment put the same point more bluntly: "Do not include in your comment text or attachments any personal identifying information or business information that you do not want published online." EPA's own docket guidance says personal information in the body of a submission gets posted to Regulations.gov, with only narrow exceptions for things like profanity or properly-flagged confidential business information. None of the three treats this as unusual — it's the baseline. A comment, attachment included, goes up as submitted; nobody downstream is editing it for privacy first.

What none of that guidance mentions

Every one of those warnings is written about content — the sentences someone typed, the figures in a table, a name signed at the bottom of a letter. None of them mentions the file's own properties: the Author field a word processor fills in without being asked, the Producer or Company name a document inherited from install settings, the timestamps a file keeps every time it's saved. Nothing in the process those agencies describe is set up to look for those fields before an attachment goes onto a public docket — "posted without change" and "as part of an attachment" both point the same direction, whether or not the guidance was written with metadata specifically in mind. A commenter who's careful about every sentence they typed can still upload a Word document whose Author field names the consulting firm that ghostwrote it, or a PDF whose Creator or Producer field carries a personal name a business filer never meant to put on the public record.

A PDF or Word file's hidden properties — before and after NearScrub ✕ Before — still in the file Author: Jane Doe Company: Acme Consulting Producer: Microsoft Word CreationDate: 2026-08-02 NearScrub in-browser, before upload ✓ After — blank before upload Author: — Company: — Producer: — CreationDate: —
The Author, Company, Producer, and timestamp fields in a PDF's Info dictionary or a Word file's docProps aren't part of what you typed — but FWS, EPA, and CFTC's own comment-posting guidance all confirm the attachment goes to the public docket exactly as uploaded, whichever column it's in when that happens.

What's actually sitting in those fields

A PDF keeps an Info dictionary — Title, Author, Subject, Keywords, Creator, Producer, and two timestamps, CreationDate and ModDate — filled in by whatever software exported the file, plus, separately, an XMP metadata stream that can duplicate or extend the same fields. A Word, Excel, or PowerPoint file — the exact three formats FWS names by name — keeps its own version split across a zip archive's docProps folder: core.xml carries the author name and the last person who saved the file, app.xml carries the Company and Manager fields (often auto-filled from the Office installation's registered organization at setup), and custom.xml carries whatever free-form properties a document-assembly or compliance tool wrote in along the way. None of it shows up in the document body. All of it travels with the file.

What NearScrub clears before the file leaves the device

NearScrub's PDF handling deletes all eight Info-dictionary keys and removes the XMP metadata stream when one is present; its Office handling blanks docProps/core.xml, app.xml, and custom.xml across docx, xlsx, pptx, and their OpenDocument equivalents. All of it runs in the browser tab itself — nothing is uploaded anywhere to perform the scrub, which is a reasonable thing to want before uploading the result to a government server for permanent public posting. It's worth being precise about what that does and doesn't reach: it clears property fields, not the document's visible text, so a name typed into a letter's signature block is exactly as present afterward as it should be for a signed comment. It doesn't resolve Word's tracked changes or delete its comments either — those live in a different part of the same zip archive and need to be accepted or removed in the word processor first. And it only reads the formats above; an older .doc or .rtf attachment isn't something NearScrub can process at all. What it's built for is the specific gap the agencies' own guidance leaves open: the fields nobody typed and nobody asked the file to fill in, that get published anyway because nothing else in the process checks for them.

Sponsored
← NearScrub

This page shows ads only if you consent.