Azure Blob Storage Unstructured File Sample Tests

Prev Next

What this pack contains

190 standard tests that run against blobs which are not tables: logs, plain text and markdown, PDFs, and images in a blob container. They run on a data source built from the Azure Blob Storage Data Source Template, and every test is a short Python script that reads the blob through the template's storage layer and leaves a data frame for Validatar to judge.

The tests are the same 190 that ship in AWS S3 Unstructured File Sample Tests. The two templates share one file-reading library and expose the same storage interface, so a script written for one runs on the other without change. Only the data source differs.

Family Tests What it covers
Text, Log and Markdown 66 Required and forbidden content (run markers, ERROR lines, PII, leaked credentials, unrendered template tokens), structure (line format, monotonic timestamps, JSONL validity, heading hierarchy, code fences, front matter, links), encoding and bytes (UTF-8, BOM, line endings, control characters, NUL bytes, mojibake), counts and budgets, rotation completeness and freshness, drift of the log-level mix
PDF 62 Opens and is not encrypted, page count and page geometry, blank pages, a classification footer on every page, DRAFT and placeholder text, personal identifiers, invoice ids reconciled to a CSV register, invoice arithmetic, printed page numbers, dates inside the reporting period, document metadata and language, embedded image count and resolution, font embedding and minimum point size, form fields, OCR on scanned pages, comparison with the previous edition
Images 62 Decode, format versus extension, colour mode, dimensions and aspect ratio, DPI, EXIF tags required and forbidden (GPS, serial numbers), copyright text, brand colour presence, palette compliance, dominant colour, white borders, colour casts, exposure, contrast, alpha coverage, blank renders, sharpness, noise, duplicates, logo detection, OCR of receipts and banners, thumbnail sets, histogram drift

Each family sits in its own subfolder under Samples - Azure Blob Unstructured Files and has its own job, so you can schedule one family without the others:

Job Runs
[Samples] Azure Blob Unstructured - Text the Text, Log and Markdown subfolder
[Samples] Azure Blob Unstructured - PDF the PDF subfolder
[Samples] Azure Blob Unstructured - Images the Images subfolder
[Samples] Azure Blob Unstructured - ALL the three jobs above, as three Job steps

The four flavours

Every family carries the same four kinds of check, so a pattern you learn on a log file transfers to a PDF or a photo.

Something must exist. The test searches the blob for content it has to contain and counts what is missing. Log - Required run markers missing looks for the start marker, the completion marker and a rows_loaded= figure. PDF - Classification footer on every page counts pages without the footer. Image - Photos missing the brand colour opens every product photo and checks that at least one percent of its pixels sit within tolerance of the brand RGB value.

Something must not exist. The test searches for content that should never be there. Text - PII patterns: SSN, card number, e-mail scans a customer notice. PDF - No DRAFT, placeholder or filler text reads every page. Image - Images carrying forbidden EXIF tags flags a GPS block, a camera serial number or an owner name in any image in the folder.

Count something. The test returns how many times a pattern occurs and compares it with an exact value or a budget. Markdown - Second-level section count expects exactly six ## headings. Log - WARN lines within budget allows up to 25. PDF - Appendix lists the expected number of distinct invoice ids expects twelve. Image - Swatch sheet shows exactly the approved number of colours quantises pixels and counts the distinct colours that cover at least half a percent of the image.

The distribution has not drifted. The test returns a table of shares that sum to 100 and compares each key with the average of the test's own last ten runs. Log - Level mix vs recent runs returns the percentage of INFO, WARN and DEBUG lines. PDF - Distribution of text across pages vs recent runs returns each page's share of the words. Image - Luminance distribution of the KPI chart vs recent runs returns an eight-bin histogram. A key that moves more than the tolerance, in percentage points, fails. The first run has no history and passes by design; it seeds the window.

Count tests come with a detail twin

Every count test has a companion whose name ends in (detail). Instead of a number, the twin returns one row per issue with two key columns, LOCATION and DETAIL, and a numeric VALUE. LOCATION is a line number, a page, a blob name or a pixel box; DETAIL is the matched text or the measurement. Log - Leaked credentials: AWS keys, passwords, tokens (detail) lists each match with its line number, the pattern name and a masked copy of the value, so the finding can be triaged without re-exposing the secret. The count version tells you how bad; the detail version tells you where.

Two further shapes appear where they fit better than a count. List tests compare what the blob contains with an expected list: Markdown - Sections match the template compares the headings found with the six the template defines, and PDF - Appendix invoice ids match the invoice register joins the ids found in the PDF to a CSV in the same container. String tests compare a single extracted value with expected text: PDF - Title metadata names the report series, Image - OCR text of the banner equals the approved copy.

Where to get it

This pack is published to the Validatar marketplace: Azure Blob Storage Unstructured File Sample Tests.

Prerequisites

  • An Azure Blob Storage data source created from the Azure Blob Storage Data Source Template, version 1.1.0 or later. Tests that list a folder recursively (INCLUDE_SUBFOLDERS = True) depend on the folder-scoping fix in 1.1.0; on 1.0.0 a listing of logs also returns logs-failing/.
  • A read-only SAS token scoped to the container is enough. The tests only list and read blobs.
  • The data source runs on a Data Agent whose Python has azure-storage-blob, pandas and numpy (the template needs these anyway), plus PyMuPDF for the PDF tests and Pillow for the image tests.
  • For the OCR tests (seven in the PDF family, eleven in the Images family): Tesseract OCR installed on the agent host and the pytesseract package in the agent's Python. Verified on a Windows agent with Tesseract 5.3.1 and the English language pack. Without it those tests error; everything else runs.
  • No metadata ingestion is required. The scripts read blobs by name and never consult the catalog.
  • Blobs to test. The pack ships pointing at the fixture described below; put a copy of the fixture in your container, or retarget the tests at your own blobs.

Import

Marketplace content imports directly. Validatar stages the pack and asks where it should land.

  1. In Validatar, open Marketplace from the main navigation.
  2. Find Azure Blob Storage - Unstructured File Sample Tests and choose Import.
  3. Pick the target project. Validatar opens that project's import screen with the pack loaded.
  4. Map the pack's data source, [Factory] Azure Blob Storage (SAS), to your Azure Blob Storage data source.
  5. Choose a target folder and commit.

The import creates the parent folder, the three subfolders, the 190 tests and the four jobs. Run the ALL job once to seed history for the drift tests, then a second time to see them compare.

To import from a downloaded file instead, take the pack XML from the marketplace listing, open Import in the target project, select the file, map the data source and commit.

Adjusting a test

Every script opens with a PARAMETERS block. Everything a user would change lives there, each value with a one-line comment, and nothing below the block needs reading to retarget the test.

# ============================ PARAMETERS ============================
# storage prefix holding the file ('./' is the bucket or container root)
FOLDER = './logs'
# object name inside FOLDER
FILE = 'etl_run_2026-09-10.log'
# name -> regular expression; every match is an issue
SECRET_PATTERNS = {'aws_access_key': '\\bAKIA[0-9A-Z]{16}\\b', 'password_assignment': '(?i)\\bpassword\\s*[=:]\\s*\\S+', ...}
#
# Fails on: FOLDER = './logs-failing', FILE = 'etl_run_secrets.log'
# Set USE_FAILING_EXAMPLE = True to run this test against that failing example.
USE_FAILING_EXAMPLE = False
if USE_FAILING_EXAMPLE:
    FOLDER = './logs-failing'
    FILE = 'etl_run_secrets.log'
# ====================================================================

Point it at your blob. Change FOLDER and FILE. FOLDER is the virtual directory written ./prefix, with ./ for the container root. Folder-scoped tests take a PATTERN glob instead of a file and, where recursion matters, an INCLUDE_SUBFOLDERS flag.

Change what it looks for. Patterns are regular expressions, palettes are lists of [R, G, B], tolerances are numbers in the unit the comment states. Edit them in the block.

See it fail. The Fails on: line names the fixture file that breaks the test, and the switch under it points the test there. Set USE_FAILING_EXAMPLE = True, run once, and you see exactly what a failure looks like before you point the test at production data. A leaked-credentials test, for example, fails on logs-failing/etl_run_secrets.log, which carries an AWS access key and a password assignment. Set the switch back to False to return to the good example.

See the file. Every test attaches the file it examined to its result with validatar_files.add(name, media_type, bytes), so the file appears in the Validatar UI next to the result. Images preview inline; PDFs, logs and text files are offered for download, and the PDF tests also attach a rendering of the first page as an image so the document can be seen without downloading it. Folder-scoped tests attach the files that raised issues first, then the first files scanned, up to the PREVIEW_FILES parameter.

Thresholds that are the pass/fail rule live in the Control data set, not in the script: the exact count a test expects, the upper bound of a budget, the expected string, the drift tolerance. Change them on the test's Control tab. The PARAMETERS block says so where it applies.

Detail twins need no extra configuration. Their control is an empty table, so every row the script returns is a failure. Rows that stop appearing stop failing.

The fixture

The pack points at a generated set of 160 blobs, the same set the S3 pack uses under identical names. Good files sit under five prefixes and every failing counterpart sits under the matching -failing prefix, so folder-scoped tests over a good prefix stay green while each failing file is one edit away. The whole set is available as a marketplace download, Unstructured File Sample Fixture (S3 and Azure Blob): unzip it and upload the six top-level folders to the root of your container, paths intact, and the pack runs as shipped.

Prefix Contents
logs/ three days of ETL logs, a JSON-lines API log, a web access log
documents/ release notes and a README in markdown, a customer notice, a fixed-width SLA report, two quarterly reports, a scanned contract, a filled form, a brochure
documents/invoices/ three one-page invoices and the CSV register they reconcile to
images/ product photos with EXIF, a brand logo, a colour swatch sheet, KPI charts, receipts and a banner for OCR, icons, a WebP hero, a grayscale TIFF scan, an animated GIF
images/thumbnails/ one thumbnail per product photo
logs-failing/, documents-failing/, images-failing/ one counterpart per failure mode: a log with ERROR lines, a notice with a Social Security number, a PDF with a DRAFT stamp, a photo carrying GPS data, a chart rendered blank, and so on

The files are generated by a script, so they carry no real data. To use the pack as shipped, place a copy of the fixture in your container under the same names. To use it on your own data, retarget each test's FOLDER and FILE and keep the Fails on: line as a record of the failure case the test was built against.

Troubleshooting

Import reports an unknown data source. Map [Factory] Azure Blob Storage (SAS) to your Azure Blob Storage data source during import. The mapping step is required.

Tests error with ResourceNotFoundError or BlobNotFound. The blob the PARAMETERS block names is not in the container the data source points at. Either upload the fixture or retarget FOLDER and FILE.

Tests error with an authorization failure. The SAS token has expired or lacks read and list permission. Re-issue it with rl permissions on the container.

OCR tests error with tesseract is not installed or it's not in your PATH. Install Tesseract on the Data Agent host, install pytesseract in the agent's Python, and restart the agent.

PDF tests error with No module named 'fitz', image tests with No module named 'PIL'. Install PyMuPDF and Pillow in the agent's Python.

A folder test counts blobs it should not. A test with INCLUDE_SUBFOLDERS = True on template 1.0.0 also lists neighbouring prefixes such as logs-failing/. Upgrade the template to 1.1.0 and re-select it on the data source, as described in the template article.

Drift tests pass on the first run no matter what. By design. Their control is the average of the test's own previous runs, and Control Missing Result Action is set to Pass so the first run seeds history. From the second run they compare.

A detail test shows rows but the count twin passes. The two read the same blob with the same parameters; check that you changed both when retargeting.

Related