
PDF tooling for Go and the command line.
View the Project on GitHub pdfcpu/pdfcpu
Use these commands to assess a directory tree of PDFs. Run them from the corpus root and record the pdfcpu version along with the results:
pdfcpu version
Validate all PDF files below the current directory:
pdfcpu validate -q "**/*.pdf"
The double quotes are important. They pass the recursive pattern to pdfcpu instead of asking the shell to expand it. This avoids operating-system command-line length limits for large corpora.
Add progress output when a file may take a long time to process:
pdfcpu validate -q --progress "**/*.pdf"
pdfcpu prints each filename before processing it, so the last progress line identifies the active file if a run stalls. Validation errors are still reported immediately.
Write progress and validation errors to a log file:
pdfcpu validate -q --progress "**/*.pdf" 2> validation.log
This redirection works in common Unix shells, PowerShell, and Windows Command Prompt. It replaces an existing log file.
Relaxed validation is the default and is suitable for an initial corpus assessment. Use strict validation for a more standards-focused pass:
pdfcpu validate -q --progress -m strict "**/*.pdf"
To assess resource optimization as well as validity, add --opt. This performs additional work and takes longer:
pdfcpu validate -q --progress --opt "**/*.pdf"
Rerun a suspect file without quiet mode to see the full result:
pdfcpu validate -m strict path/to/problem.pdf
For a corpus run, pdfcpu reports each invalid file and finishes with a summary such as:
validation failed: 66 of 463 files invalid
The process exits with status 0 only when every selected PDF validates. Check the status immediately after the
command using the syntax for your shell:
| Shell | Command |
|---|---|
| Bash or zsh | echo $? |
| PowerShell | echo $LASTEXITCODE |
| Windows Command Prompt | echo %ERRORLEVEL% |
The file count covers inputs matching *.pdf, not every file in the directory tree. Recursive input is processed in
pdfcpu’s stable lexical order, which may differ from Finder or another file manager’s display order.
A failed corpus run is only a starting point.
Before opening an issue, reproduce one failure individually with the
latest release or prerelease and isolate one problem with the smallest shareable PDF.
Include the pdfcpu version, operating system, exact command, final summary, and the relevant error message. Attach large logs as compressed files instead of pasting them, and remove sensitive filenames and paths.
Corpus output is diagnostic input, not an issue backlog.
Do not use AI, bots, or scripts to create issues in bulk or one issue per validation error.
AI-generated, bulk-generated, or mechanically reformatted corpus reports will be closed immediately without
investigation.
Repeated bulk submissions may be treated as spam.
Every submitted report must be understood, manually verified, and independently reproduced by its author.
Public issues are for isolated, reproducible pdfcpu problems.
Reviewing an entire corpus, classifying and
prioritizing large result sets, investigating confidential or non-shareable files, and delivering fixes to an agreed
schedule are substantial engineering work and may require a paid engagement.
Limited corpus investigation and remediation work may be available by arrangement.
See Validate for the complete command reference.