What this pack contains
190 standard tests that run against files which are not tables: logs, plain text and markdown, PDFs, and images in a folder tree on the Data Agent host, on local disk or a UNC share. They run on a data source built from the Windows Directory Data Source Template, and every test is a short Python script that reads the file through the template's storage layer and leaves a data frame for Validatar to judge. The tests are the same 190 that ship in the AWS S3 and Azure Blob Storage packs; only the data source differs.
| Family | Tests | What it covers |
|---|---|---|
| Text, Log and Markdown | 66 | Required and forbidden content (run markers, ERROR lines, PII, leaked credentials, unrendered template tokens), structure (line format, monotonic timestamps, JSONL validity, heading hierarchy, code fences, front matter, links), encoding and bytes (UTF-8, BOM, line endings, control characters, NUL bytes, mojibake), counts and budgets, rotation completeness and freshness, drift of the log-level mix |
| 62 | Opens and is not encrypted, page count and page geometry, blank pages, a classification footer on every page, DRAFT and placeholder text, personal identifiers, invoice ids reconciled to a CSV register, invoice arithmetic, printed page numbers, dates inside the reporting period, document metadata and language, embedded image count and resolution, font embedding and minimum point size, form fields, OCR on scanned pages, comparison with the previous edition | |
| Images | 62 | Decode, format versus extension, colour mode, dimensions and aspect ratio, DPI, EXIF tags required and forbidden (GPS, serial numbers), copyright text, brand colour presence, palette compliance, dominant colour, white borders, colour casts, exposure, contrast, alpha coverage, blank renders, sharpness, noise, duplicates, logo detection, OCR of receipts and banners, thumbnail sets, histogram drift |
Each family sits in its own subfolder under Samples - Windows Directory Unstructured Files and has its own job, so you can schedule one family without the others:
| Job | Runs |
|---|---|
[Samples] Windows Directory Unstructured - Text |
the Text, Log and Markdown subfolder |
[Samples] Windows Directory Unstructured - PDF |
the PDF subfolder |
[Samples] Windows Directory Unstructured - Images |
the Images subfolder |
[Samples] Windows Directory Unstructured - ALL |
the three jobs above, as three Job steps |
The four flavours
Every family carries the same four kinds of check, so a pattern you learn on a log file transfers to a PDF or a photo.
Something must exist. The test searches the file for content it has to contain and counts what is missing. Log - Required run markers missing looks for the start marker, the completion marker and a rows_loaded= figure. PDF - Classification footer on every page counts pages without the footer. Image - Photos missing the brand colour opens every product photo and checks that at least one percent of its pixels sit within tolerance of the brand RGB value.
Something must not exist. The test searches for content that should never be there. Text - PII patterns: SSN, card number, e-mail scans a customer notice. PDF - No DRAFT, placeholder or filler text reads every page. Image - Images carrying forbidden EXIF tags flags a GPS block, a camera serial number or an owner name in any image in the folder.
Count something. The test returns how many times a pattern occurs and compares it with an exact value or a budget. Markdown - Second-level section count expects exactly six ## headings. Log - WARN lines within budget allows up to 25. PDF - Appendix lists the expected number of distinct invoice ids expects twelve. Image - Swatch sheet shows exactly the approved number of colours quantises pixels and counts the distinct colours that cover at least half a percent of the image.
The distribution has not drifted. The test returns a table of shares that sum to 100 and compares each key with the average of the test's own last ten runs. Log - Level mix vs recent runs returns the percentage of INFO, WARN and DEBUG lines. PDF - Distribution of text across pages vs recent runs returns each page's share of the words. Image - Luminance distribution of the KPI chart vs recent runs returns an eight-bin histogram. A key that moves more than the tolerance, in percentage points, fails. The first run has no history and passes by design; it seeds the window.
Count tests come with a detail twin
Every count test has a companion whose name ends in (detail). Instead of a number, the twin returns one row per issue with two key columns, LOCATION and DETAIL, and a numeric VALUE. LOCATION is a line number, a page, a file name or a pixel box; DETAIL is the matched text or the measurement. Log - Leaked credentials: AWS keys, passwords, tokens (detail) lists each match with its line number, the pattern name and a masked copy of the value, so the finding can be triaged without re-exposing the secret. The count version tells you how bad; the detail version tells you where.
Two further shapes appear where they fit better than a count. List tests compare what the file contains with an expected list: Markdown - Sections match the template compares the headings found with the six the template defines, and PDF - Appendix invoice ids match the invoice register joins the ids found in the PDF to a CSV in a neighbouring folder. String tests compare a single extracted value with expected text: PDF - Title metadata names the report series, Image - OCR text of the banner equals the approved copy.
Where to get it
This pack is published to the Validatar marketplace: Windows Directory Unstructured File Sample Tests.
Prerequisites
- A data source created from the Windows Directory Data Source Template, version 2.1.0 or later.
- A Data Agent. The scripts run on the agent host, which is where the folder has to be readable: local disk, or a UNC share the agent's service account can open. The agent's Python needs
pandasandnumpy(the template needs these anyway), plusPyMuPDFfor the PDF tests andPillowfor the image tests. - For the OCR tests (seven in the PDF family, eleven in the Images family): Tesseract OCR installed on the agent host and the
pytesseractpackage in the agent's Python. Verified on a Windows agent with Tesseract 5.3.1 and the English language pack. Without it those tests error; everything else runs. - No metadata ingestion is required. The scripts read files by path and never consult the catalog.
- Files to test. The pack ships pointing at the fixture described below; copy the fixture under the data source's Parent Folder Path, or retarget the tests at your own files.
Import
Marketplace content imports directly. Validatar stages the pack and asks where it should land.
- In Validatar, open Marketplace from the main navigation.
- Find Windows Directory - Unstructured File Sample Tests and choose Import.
- Pick the target project. Validatar opens that project's import screen with the pack loaded.
- Map the pack's data source, [Factory] Windows Directory (Local Agent), to your Windows Directory data source.
- Choose a target folder and commit.
The import creates the parent folder, the three subfolders, the 190 tests and the four jobs. Run the ALL job once to seed history for the drift tests, then a second time to see them compare.
To import from a downloaded file instead, take the pack XML from the marketplace listing, open Import in the target project, select the file, map the data source and commit.
Adjusting a test
Every script opens with a PARAMETERS block. Everything a user would change lives there, each value with a one-line comment, and nothing below the block needs reading to retarget the test.
# ============================ PARAMETERS ============================
# storage prefix holding the file ('./' is the bucket or container root)
FOLDER = './logs'
# object name inside FOLDER
FILE = 'etl_run_2026-09-10.log'
# name -> regular expression; every match is an issue
SECRET_PATTERNS = {'aws_access_key': '\\bAKIA[0-9A-Z]{16}\\b', 'password_assignment': '(?i)\\bpassword\\s*[=:]\\s*\\S+', ...}
#
# Fails on: FOLDER = './logs-failing', FILE = 'etl_run_secrets.log'
# Set USE_FAILING_EXAMPLE = True to run this test against that failing example.
USE_FAILING_EXAMPLE = False
if USE_FAILING_EXAMPLE:
FOLDER = './logs-failing'
FILE = 'etl_run_secrets.log'
# ====================================================================
Point it at your file. Change FOLDER and FILE. FOLDER is a folder relative to the data source's Parent Folder Path, written ./folder, with ./ for the root; a subfolder is ./documents/invoices. The comment in the block mentions buckets and containers because the same scripts ship in the S3 and Azure packs, and on a Windows data source the root is the Parent Folder Path. Folder-scoped tests take a PATTERN glob instead of a file and, where recursion matters, an INCLUDE_SUBFOLDERS flag.
Change what it looks for. Patterns are regular expressions, palettes are lists of [R, G, B], tolerances are numbers in the unit the comment states. Edit them in the block.
See it fail. The Fails on: line names the fixture file that breaks the test, and the switch under it points the test there. Set USE_FAILING_EXAMPLE = True, run once, and you see exactly what a failure looks like before you point the test at production data. A leaked-credentials test, for example, fails on logs-failing\etl_run_secrets.log, which carries an AWS access key and a password assignment. Set the switch back to False to return to the good example.
See the file. Every test attaches the file it examined to its result with validatar_files.add(name, media_type, bytes), so the file appears in the Validatar UI next to the result. Images preview inline; PDFs, logs and text files are offered for download, and the PDF tests also attach a rendering of the first page as an image so the document can be seen without downloading it. Folder-scoped tests attach the files that raised issues first, then the first files scanned, up to the PREVIEW_FILES parameter.
Thresholds that are the pass/fail rule live in the Control data set, not in the script: the exact count a test expects, the upper bound of a budget, the expected string, the drift tolerance. Change them on the test's Control tab. The PARAMETERS block says so where it applies.
Detail twins need no extra configuration. Their control is an empty table, so every row the script returns is a failure. Rows that stop appearing stop failing.
The fixture
The pack points at a generated set of 160 files. Good files sit under five folders and every failing counterpart sits under the matching -failing folder, so folder-scoped tests over a good folder stay green while each failing file is one edit away.
| Folder | Contents |
|---|---|
logs\ |
three days of ETL logs, a JSON-lines API log, a web access log |
documents\ |
release notes and a README in markdown, a customer notice, a fixed-width SLA report, two quarterly reports, a scanned contract, a filled form, a brochure |
documents\invoices\ |
three one-page invoices and the CSV register they reconcile to |
images\ |
product photos with EXIF, a brand logo, a colour swatch sheet, KPI charts, receipts and a banner for OCR, icons, a WebP hero, a grayscale TIFF scan, an animated GIF |
images\thumbnails\ |
one thumbnail per product photo |
logs-failing\, documents-failing\, images-failing\ |
one counterpart per failure mode: a log with ERROR lines, a notice with a Social Security number, a PDF with a DRAFT stamp, a photo carrying GPS data, a chart rendered blank, and so on |
The files are generated by a script, so they carry no real data. The whole set is available as a marketplace download, Unstructured File Sample Fixture (S3 and Azure Blob). Unzip it on the Data Agent host and copy the six top-level folders under the data source's Parent Folder Path, structure intact, and the pack runs as shipped. The zip's own README describes S3 and Azure uploads; for a Windows data source the copy is the whole step. To use the pack on your own data, retarget each test's FOLDER and FILE and keep the Fails on: line as a record of the failure case the test was built against.
Troubleshooting
Import reports an unknown data source. Map [Factory] Windows Directory (Local Agent) to your Windows Directory data source during import. The mapping step is required.
Tests error with FileNotFoundError. The file the PARAMETERS block names is not under the Parent Folder Path the data source points at. Either copy the fixture there or retarget FOLDER and FILE.
Tests error with PermissionError, or a folder test finds no files. The Data Agent's service account cannot read the folder. Grant it read permission on the folder and, for a UNC path, on the share as well. This is the same condition the template article describes for an empty catalog.
OCR tests error with tesseract is not installed or it's not in your PATH. Install Tesseract on the Data Agent host, install pytesseract in the agent's Python, and restart the agent.
PDF tests error with No module named 'fitz', image tests with No module named 'PIL'. Install PyMuPDF and Pillow in the agent's Python.
Drift tests pass on the first run no matter what. By design. Their control is the average of the test's own previous runs, and Control Missing Result Action is set to Pass so the first run seeds history. From the second run they compare.
A detail test shows rows but the count twin passes. The two read the same file with the same parameters; check that you changed both when retargeting.
Related
- Windows Directory Data Source Template
- AWS S3 Unstructured File Sample Tests, the same 190 tests against an S3 bucket
- Azure Blob Storage Unstructured File Sample Tests, the same 190 tests against a blob container
- Macros
- Python Script Data Source Template