What this pack contains
190 standard tests that run against files which are not tables: logs, plain text and markdown, PDFs, and images stored in S3. They run on a data source built from the AWS S3 Data Source Template, and every test is a short Python script that reads the file through the template's storage layer and leaves a data frame for Validatar to judge.
| Family | Tests | What it covers |
|---|---|---|
| Text, Log and Markdown | 66 | Required and forbidden content (run markers, ERROR lines, PII, leaked credentials, unrendered template tokens), structure (line format, monotonic timestamps, JSONL validity, heading hierarchy, code fences, front matter, links), encoding and bytes (UTF-8, BOM, line endings, control characters, NUL bytes, mojibake), counts and budgets, rotation completeness and freshness, drift of the log-level mix |
| 62 | Opens and is not encrypted, page count and page geometry, blank pages, a classification footer on every page, DRAFT and placeholder text, personal identifiers, invoice ids reconciled to a CSV register, invoice arithmetic, printed page numbers, dates inside the reporting period, document metadata and language, embedded image count and resolution, font embedding and minimum point size, form fields, OCR on scanned pages, comparison with the previous edition | |
| Images | 62 | Decode, format versus extension, colour mode, dimensions and aspect ratio, DPI, EXIF tags required and forbidden (GPS, serial numbers), copyright text, brand colour presence, palette compliance, dominant colour, white borders, colour casts, exposure, contrast, alpha coverage, blank renders, sharpness, noise, duplicates, logo detection, OCR of receipts and banners, thumbnail sets, histogram drift |
Each family sits in its own subfolder under Samples - S3 Unstructured Files and has its own job, so you can schedule one family without the others:
| Job | Runs |
|---|---|
[Samples] S3 Unstructured - Text |
the Text, Log and Markdown subfolder |
[Samples] S3 Unstructured - PDF |
the PDF subfolder |
[Samples] S3 Unstructured - Images |
the Images subfolder |
[Samples] S3 Unstructured - ALL |
the three jobs above, as three Job steps |
The four flavours
Every family carries the same four kinds of check, so a pattern you learn on a log file transfers to a PDF or a photo.
Something must exist. The test searches the file for content it has to contain and counts what is missing. Log - Required run markers missing looks for the start marker, the completion marker and a rows_loaded= figure. PDF - Classification footer on every page counts pages without the footer. Image - Photos missing the brand colour opens every product photo and checks that at least one percent of its pixels sit within tolerance of the brand RGB value.
Something must not exist. The test searches for content that should never be there. Text - PII patterns: SSN, card number, e-mail scans a customer notice. PDF - No DRAFT, placeholder or filler text reads every page. Image - Images carrying forbidden EXIF tags flags a GPS block, a camera serial number or an owner name in any image in the folder.
Count something. The test returns how many times a pattern occurs and compares it with an exact value or a budget. Markdown - Second-level section count expects exactly six ## headings. Log - WARN lines within budget allows up to 25. PDF - Appendix lists the expected number of distinct invoice ids expects twelve. Image - Swatch sheet shows exactly the approved number of colours quantises pixels and counts the distinct colours that cover at least half a percent of the image.
The distribution has not drifted. The test returns a table of shares that sum to 100 and compares each key with the average of the test's own last ten runs. Log - Level mix vs recent runs returns the percentage of INFO, WARN and DEBUG lines. PDF - Distribution of text across pages vs recent runs returns each page's share of the words. Image - Luminance distribution of the KPI chart vs recent runs returns an eight-bin histogram. A key that moves more than the tolerance, in percentage points, fails. The first run has no history and passes by design; it seeds the window.
Count tests come with a detail twin
Every count test has a companion whose name ends in (detail). Instead of a number, the twin returns one row per issue with two key columns, LOCATION and DETAIL, and a numeric VALUE. LOCATION is a line number, a page, a file name or a pixel box; DETAIL is the matched text or the measurement. Log - Leaked credentials: AWS keys, passwords, tokens (detail) lists each match with its line number, the pattern name and a masked copy of the value, so the finding can be triaged without re-exposing the secret. The count version tells you how bad; the detail version tells you where.
Two further shapes appear where they fit better than a count. List tests compare what the file contains with an expected list: Markdown - Sections match the template compares the headings found with the six the template defines, and PDF - Appendix invoice ids match the invoice register joins the ids found in the PDF to a CSV in the same bucket. String tests compare a single extracted value with expected text: PDF - Title metadata names the report series, Image - OCR text of the banner equals the approved copy.
Where to get it
This pack is published to the Validatar marketplace: AWS S3 Unstructured File Sample Tests.
Prerequisites
- An S3 data source created from the AWS S3 Data Source Template, version 1.1.1 or later. Tests that list a folder recursively (
INCLUDE_SUBFOLDERS = True) depend on the folder-scoping fix in 1.1.1; on 1.1.0 a listing oflogsalso returnslogs-failing/. - The data source runs on a Data Agent whose Python has
boto3,pandasandnumpy(the template needs these anyway), plusPyMuPDFfor the PDF tests andPillowfor the image tests. - For the OCR tests (seven in the PDF family, eleven in the Images family): Tesseract OCR installed on the agent host and the
pytesseractpackage in the agent's Python. Verified on a Windows agent with Tesseract 5.3.1 and the English language pack. Without it those tests error; everything else runs. - No metadata ingestion is required. The scripts read objects by key and never consult the catalog.
- Files to test. The pack ships pointing at the fixture described below; put a copy of the fixture in your bucket, or retarget the tests at your own files.
Import
Marketplace content imports directly. Validatar stages the pack and asks where it should land.
- In Validatar, open Marketplace from the main navigation.
- Find AWS S3 - Unstructured File Sample Tests and choose Import.
- Pick the target project. Validatar opens that project's import screen with the pack loaded.
- Map the pack's data source, [Factory] AWS S3 (Multi-Format + JSON), to your S3 data source.
- Choose a target folder and commit.
The import creates the parent folder, the three subfolders, the 190 tests and the four jobs. Run the ALL job once to seed history for the drift tests, then a second time to see them compare.
To import from a downloaded file instead, take the pack XML from the marketplace listing, open Import in the target project, select the file, map the data source and commit.
Adjusting a test
Every script opens with a PARAMETERS block. Everything a user would change lives there, each value with a one-line comment, and nothing below the block needs reading to retarget the test.
# ============================ PARAMETERS ============================
# storage prefix holding the file ('./' is the bucket or container root)
FOLDER = './logs'
# object name inside FOLDER
FILE = 'etl_run_2026-09-10.log'
# name -> regular expression; every match is an issue
SECRET_PATTERNS = {'aws_access_key': '\\bAKIA[0-9A-Z]{16}\\b', 'password_assignment': '(?i)\\bpassword\\s*[=:]\\s*\\S+', ...}
#
# Fails on: FOLDER = './logs-failing', FILE = 'etl_run_secrets.log'
# Set USE_FAILING_EXAMPLE = True to run this test against that failing example.
USE_FAILING_EXAMPLE = False
if USE_FAILING_EXAMPLE:
FOLDER = './logs-failing'
FILE = 'etl_run_secrets.log'
# ====================================================================
Point it at your file. Change FOLDER and FILE. FOLDER is the key prefix written ./prefix, with ./ for the bucket root. Folder-scoped tests take a PATTERN glob instead of a file and, where recursion matters, an INCLUDE_SUBFOLDERS flag.
Change what it looks for. Patterns are regular expressions, palettes are lists of [R, G, B], tolerances are numbers in the unit the comment states. Edit them in the block.
See it fail. The Fails on: line names the fixture file that breaks the test, and the switch under it points the test there. Set USE_FAILING_EXAMPLE = True, run once, and you see exactly what a failure looks like before you point the test at production data. A leaked-credentials test, for example, fails on logs-failing/etl_run_secrets.log, which carries an AWS access key and a password assignment. Set the switch back to False to return to the good example.
See the file. Every test attaches the file it examined to its result with validatar_files.add(name, media_type, bytes), so the file appears in the Validatar UI next to the result. Images preview inline; PDFs, logs and text files are offered for download, and the PDF tests also attach a rendering of the first page as an image so the document can be seen without downloading it. Folder-scoped tests attach the files that raised issues first, then the first files scanned, up to the PREVIEW_FILES parameter.
Thresholds that are the pass/fail rule live in the Control data set, not in the script: the exact count a test expects, the upper bound of a budget, the expected string, the drift tolerance. Change them on the test's Control tab. The PARAMETERS block says so where it applies.
Detail twins need no extra configuration. Their control is an empty table, so every row the script returns is a failure. Rows that stop appearing stop failing.
The fixture
The pack points at a generated set of 160 objects. Good files sit under five prefixes and every failing counterpart sits under the matching -failing prefix, so folder-scoped tests over a good prefix stay green while each failing file is one edit away.
| Prefix | Contents |
|---|---|
logs/ |
three days of ETL logs, a JSON-lines API log, a web access log |
documents/ |
release notes and a README in markdown, a customer notice, a fixed-width SLA report, two quarterly reports, a scanned contract, a filled form, a brochure |
documents/invoices/ |
three one-page invoices and the CSV register they reconcile to |
images/ |
product photos with EXIF, a brand logo, a colour swatch sheet, KPI charts, receipts and a banner for OCR, icons, a WebP hero, a grayscale TIFF scan, an animated GIF |
images/thumbnails/ |
one thumbnail per product photo |
logs-failing/, documents-failing/, images-failing/ |
one counterpart per failure mode: a log with ERROR lines, a notice with a Social Security number, a PDF with a DRAFT stamp, a photo carrying GPS data, a chart rendered blank, and so on |
The files are generated by a script, so they carry no real data. The whole set is available as a marketplace download, Unstructured File Sample Fixture (S3 and Azure Blob): unzip it and upload the six top-level folders to the root of your bucket, paths intact, and the pack runs as shipped. To use it on your own data, retarget each test's FOLDER and FILE and keep the Fails on: line as a record of the failure case the test was built against.
Troubleshooting
Import reports an unknown data source. Map [Factory] AWS S3 (Multi-Format + JSON) to your S3 data source during import. The mapping step is required.
Tests error with NoSuchKey. The object the PARAMETERS block names is not in the bucket the data source points at. Either upload the fixture or retarget FOLDER and FILE.
OCR tests error with tesseract is not installed or it's not in your PATH. Install Tesseract on the Data Agent host, install pytesseract in the agent's Python, and restart the agent.
PDF tests error with No module named 'fitz', image tests with No module named 'PIL'. Install PyMuPDF and Pillow in the agent's Python.
A folder test counts files it should not. A test with INCLUDE_SUBFOLDERS = True on template 1.1.0 or earlier also lists neighbouring prefixes such as logs-failing/. Upgrade the template to 1.1.1 and re-select it on the data source, as described in the template article.
Drift tests pass on the first run no matter what. By design. Their control is the average of the test's own previous runs, and Control Missing Result Action is set to Pass so the first run seeds history. From the second run they compare.
A detail test shows rows but the count twin passes. The two read the same file with the same parameters; check that you changed both when retargeting.
Related
- AWS S3 Data Source Template
- Azure Blob Storage Unstructured File Sample Tests, the same 190 tests against a blob container
- Macros
- Python Script Data Source Template