reliability job / in production
Backup Verification
A nightly job that answers one question: if we had to restore from last night's backup, would the data actually be there and readable?

the problem.
A backup job that says "succeeded" doesn't prove the files inside it are intact or that you could restore them when it matters. You usually find out on the worst possible day.
file shares.
It picks yesterday's Azure Files snapshot, downloads a sample of files and opens each one with a real parser: FoxPro tables and their memo and index files, Excel, CSV, PDF, ZIP, RTF and more. The sample rotates, so over time it covers far more than the same few files. Once a week it fully parses the big files too.
It also watches for what ransomware leaves behind: content that suddenly looks encrypted, ransom note file names, files that disappeared overnight or a large share of a folder rewritten at once. Missing daily snapshots are flagged too.
vm disks.
For each protected VM it confirms there's a recent recovery point and whether a fast restore is still possible or only the slower vault copy is left.
how it runs.
It moved from a dedicated PC under Task Scheduler to a Docker image on Azure Kubernetes Service, triggered by Airflow every night. Results go out as a branded HTML email through Microsoft Graph that works in light and dark mode, with an optional Teams alert that only speaks up when something is wrong.
how it works.
- pick yesterday's snapshot
- download a rotating sample of files
- open each file with a real parser
- check for encryption, drift and missing snapshots
- check VM recovery points
- email the report, alert Teams on failures