← Back to writeups
Personal project Windows agent · FastAPI backend · Supabase · Render

Building an EDR from the ground up taught me more than using one ever could.

SentryOS is an endpoint detection and response platform I designed, built, and now actually run: a Windows agent, a backend that decides what's worth flagging, and a console my friends log into to see their own machines, and nobody else's.

Critical T1562.001

Impair Defenses: Disable or Modify Tools

DESKTOP-T30DKRU · process_create

A PowerShell process launched cmd.exe to run "Set-MpPreference -DisableRealtimeMonitoring True," attempting to turn off Windows Defender real-time protection. First detection of this technique on this host.
powershell.exe → cmd.exe

What it actually is

SentryOS started as a school project and became something with real stakes: a Windows Service that watches for suspicious activity on a machine, a backend that decides what's worth a human's attention, and a console for reviewing and acting on what it finds. It's not a demo dressed up to look real. The telemetry is real Windows Sysmon and Security-log data, the detections fire against actual activity, and a small group of real people log in to watch their own machines.

This is the story of building all three pieces, and the decisions, trade-offs, and bugs that shaped what I ended up with.

How it fits together

Agent on the endpoint, backend deciding what matters, console for the humans

Agent
Windows Service · Sysmon + native Security log · PyInstaller
→
Backend
FastAPI on Render · rule matching · AI triage (Groq)
→
Database
Supabase (Postgres)

The console is server-rendered off the same backend, with no separate frontend service and no API layer to keep in sync with a UI.

Feature tour

The console, in the order you'd actually use it

Dashboard

A glanceable overview: severity counts, top techniques, and a real agent-health indicator that polls actual check-ins rather than showing a hardcoded "online" badge.

SentryOS dashboard showing detection counts, severity distribution, top techniques, and reporting endpoints
dashboard-sanitized.png

Before going further: how this is actually running

This isn't a local demo. The backend is a FastAPI service deployed on Render, and every event, rule, user, and license key lives in a Supabase Postgres database, both free-tier, both run like a real (if small) production service, cold starts and all. Everything below is a screenshot of that live system, not a mockup.

Alerts, with real process lineage

Click into any detection and see the actual parent → child process chain pulled from Sysmon telemetry, not a description of it. Bulk actions let you resolve, tag a verdict, or leave a note across several alerts at once.

Incidents

Related alerts on the same host, close together in time, get correlated into one incident instead of showing up as five disconnected rows, a small step toward telling a story instead of a list.

SentryOS Incidents page showing alerts on the same host correlated into one incident
incidents-sanitized.png

Endpoints

Every machine that's ever activated shows up here, not just the ones that happened to trigger an alert. Agent status (online, stale, offline), reported version, and a per-host severity mix make this the place to check whether the fleet is actually healthy, a separate question from whether anything's actually wrong.

SentryOS Endpoints page listing every registered host with agent status, version, and severity mix
endpoints-sanitized.png

AI Analyst

Pick an alert, ask it a plain question ("did this touch the registry," "is this the same process as the last one"), and get an answer grounded in that specific event's raw data. It only ever reasons about the one alert you point it at, not a general chat about the whole environment.

AI Analyst page showing a question asked about a specific alert and its grounded answer
ai-analyst-sanitized.png

Event Search

The actual data lake: every event the agent ever sent, matched by a rule or not, searchable by host, type, or free text. Alerts answers "what's worth acting on"; this answers "what actually happened," and the two are deliberately different questions with deliberately different answers.

Event Search page showing raw events searchable by host, type, and free text
event-search-sanitized.png

A rules engine anyone can use

Detection logic lives in the database, not in code: write a new rule from the console and it takes effect on the very next event. Rules can be scoped, so one person's custom detection only ever fires on their own machines.

Real account boundaries

Every user is scoped to their own set of endpoints, enforced on the server, not just hidden in the UI. A member can write their own rules and never sees, or affects, anyone else's machines.

Exceptions

An allowlist that sits above rule matching: a specific file hash or process name can get suppressed without touching the rule that caught it, so one known-good file stops alerting without weakening the rule for everything else it's supposed to catch.

SentryOS Exceptions page showing allowlisted hashes and process names
exceptions-sanitized.png

Settings

The actual admin control surface behind everything above: creating scopes, assigning endpoints to them, adding users with a role and a scope, generating license keys, and publishing the current agent version. This is where "multi-tenant" stops being a design decision and becomes a page someone actually clicks through.

SentryOS Settings page showing scope, user, and license key management
settings-sanitized.png

Audit log

Who resolved an alert, who created a rule, who added a user, and when, going back as far as the log retains. The kind of thing that turns "I think someone changed that rule" into an actual answer instead of a guess.

SentryOS Audit log page showing a history of actions by user and timestamp
audit-log-sanitized.png

Login and account recovery

Five failed attempts locks an account for fifteen minutes. There's no email service wired up, so instead of a self-serve "forgot password" form, an admin generates a single-use reset link from Settings and sends it directly, solving the same real problem without needing an email pipeline this project doesn't have.

The agent itself

Runs as an actual Windows Service: SYSTEM-level, starts at boot, no logged-in user required. It reads native Windows auditing for process creation, and Sysmon for network connections, cross-process memory access, and file writes.

agent.exe running in debug mode, terminal output showing a Sysmon event caught and forwarded to the backend
agent-terminal-sanitized.png

An installer that does the boring parts for you

One double-click: prompts for a license key, silently turns on the Windows auditing settings the agent needs, installs Sysmon with a scoped config, and registers the agent as a service, with no PowerShell and no manual steps.

Where it got interesting

The bugs and decisions that actually taught me something

The agent that was watching itself

After wiring the agent to read Sysmon's network-connection events, I noticed a wall of repeated entries: the same destination, over and over, several times a second. It was the agent's own outbound calls to the backend: Sysmon was logging the agent reporting an event, which was itself a network connection, which the agent then reported. Found it by counting raw event volume live on a test machine and tracing the flood back to a single source process. Fixed with one exclusion rule in the Sysmon config.

A fix that quietly never applied, for weeks

Every time I updated the Sysmon detection config and reinstalled, nothing changed. The actual cause: Sysmon refuses to reinstall its configuration over an already-installed copy, and the installer's silent execution swallowed that failure completely, with no error, no warning, just an installer that appeared to succeed while doing nothing. Only surfaced by dumping Sysmon's live configuration directly and noticing it didn't match what I'd just shipped.

Losing a detection to a network blip

A transient database disconnect early in the ingestion pipeline could silently drop a real detection before it was ever evaluated. The fix wasn't "retry everything." It was deciding, function by function, what's worth a retry (the actual alert) versus what's safe to log and move past (a timestamp update). Losing the wrong kind of write is a much bigger problem than losing the other.

A service that only worked when I restarted it myself

Windows won't retry a service that fails on its first start; it just stays down. Since the very first thing the agent does is activate over the network, any one-off hiccup at that exact moment could permanently kill an unattended install with nobody watching. Fixed at the OS level with a recovery policy, not by telling people to reboot when something seems off.

A login screen isn't an account boundary

Early multi-user support trusted a scope selection stored in a cookie, which meant it was trusted, and editable, entirely on the visitor's own machine. Real separation meant re-checking the account's actual permissions on every single request, server-side, and never taking the browser's word for what it's allowed to see.

What it doesn't do

Worth saying plainly, not hiding