Skip to content
Open for consulting and new roles
greeklinux, Agentic AI Security Architect

greeklinux

Agentic AI Security Architect

Open to new opportunities

Hi, I am greeklinux.

Agentic AI Security Architect

I build security-focused software, cloud automation, and AI systems. Explore the projects, the controls behind their design, and the technologies that connect them.

I explore these ideas hands-on. I run my own custom labs and build my own projects, breaking and rebuilding systems until the concepts are second nature. That self-driven work is where much of what I know was actually earned.

Certified across the full Microsoft security and AI stack: 10 current certifications, plus 2 expired and open to renewal and 3 taken or booked but not yet conferred. SC-100, SC-500, and SC-900 on security; AZ-305 and AZ-104 on cloud and architecture; and a deep AI bench in AI-900, AI-102, AI-103, AB-100, and AB-741. Security+ is the expired one, and the AI-500 Multi-Agent AI Solutions Expert beta, AI-200, and CISSP are the pending ones. If Microsoft makes it, I have probably deployed it, secured it, or automated it.

Cloud Security

Defender XDRSentinelCortex XDRSentinelOneQualysTenableKQLSOARZero TrustConditional AccessIntunePurviewDMARCDefender for CloudSecure ScoreEntra

AI / Agentic

Claude CodeSemantic KernelAzure AI FoundryRAGMCPCopilot StudioOllamaLM StudioDeepSeekKimiGeminiAntigravityHermesGrokNIST AI RMF

Blockchain

Web3 securitySmart contractsWallet securityOn-chain data pipelinesPrediction marketsMarket data APIs

Software Engineering

PythonPowerShellTypeScriptSQLDockerLinuxGitGitHub ActionsCI/CDREST APIsBash
Defender XDRSentinelIntunePurviewCloud SecurityAgentic AIThreat HuntingBlockchainZero TrustPersonalIncident ResponseCloud PlatformsAI GovernanceGREEKLINUX

Cloud SecurityDefender XDR | Sentinel | Intune | Purview

Hardening Azure and M365 at scale: ASR rules, tamper protection, CIS baselines.

0
Current certifications
0
Technologies in the toolkit

About

Who I am

My focus is the intersection of cloud security, identity, and agentic AI. I like turning complex technical ideas into tools with clear boundaries, useful interfaces, and results that can be checked.

What I care about is making systems real and durable, not just secure. I bring the same discipline to cloud security, identity, and operations that I bring to software engineering: controls and processes that are engineered, compliant, and repeatable, never one-off heroics. I build automation that is self-learning and self-healing, wrapped in security controls that stay tightly guarded, so the system improves itself and recovers on its own without ever loosening the guardrails.

In practice that means hunting with KQL mapped to MITRE ATT&CK, automating response so teams scale, hardening identity and cloud to Zero Trust, and designing AI governance so new technology gets adopted safely. I am equally at home on an incident bridge and in a strategy planning session, translating deep technical work into risk and business terms.

And it is not only for the enterprise. The same skills turn personal projects and everyday errands into automated, self-running systems, and build finance tools that help people actually earn. Whether you are a company hardening its stack or a person trying to save hours and make money, I build for both work and life.

Right now I am focused on the intersection that matters most: as AI gets more autonomy, the security of the system around it becomes the product. That is what I build.

At a glance

  • Cloud security and SOC operations
  • Agentic AI architecture and safety
  • AI solutions and cost-cutting consulting
  • Zero Trust and identity
  • AI governance for regulated fields
  • Remote, and open to relocation

Raised Microsoft Secure Score from 35% to the 80 to 85% band for multiple clients, with no outages and no major operational disruption.

Deployed the full Defender suite top to bottom: Endpoint, Identity, Cloud, Storage, and Office, rolled out through ringed pilot groups.

Architected and hardened cloud infrastructure across Azure, Microsoft 365, and AWS, with network segmentation, workload protection, and identity-aware access.

Built a Zero Trust identity program with Conditional Access, Identity Protection, self-service password reset, and token protection, and drove email to full DMARC enforcement.

Engineered a multi-provider agentic AI engine with air-gapped inference for high-sensitivity security data.

Authored the AI governance program, the policies and controls that keep AI safe, legal, and accountable, aligned to NIST AI RMF, ISO/IEC 42001, the EU AI Act, and HHS AI strategy.

Off the clock

Chess, lots of chess

Pattern recognition, calculated risk, thinking three moves ahead. It is the same muscle I use in threat hunting.

Piano

Discipline and precision with an output you can feel. The practice habit transfers to everything else I do.

Reading strategy

Robert Greene is my favorite author, The 48 Laws of Power especially. I read for how people and systems actually behave.

Building, not gaming

Surprisingly, no video games. My free compute goes into agentic AI side projects and market-data systems instead.

Selected work

Projects

In order: the two builds I own end to end, the deployment I am best known for, the case studies behind them, and then every pipeline drawn out stage by stage.

FeaturedAgentic AI, self-learning, paper-only by design

One of my three best builds, kept off the front page. Type the number below to open it.

Generating a number.

Flagship buildSelf-learning, multi-model, paper only

PolyMind: Self-Learning, Multi-Model Paper-Trading Intelligence Platform

A research platform for comparing AI forecasts across four market domains using simulated trading. Its public examples focus on evidence quality, uncertainty, learning controls, and the difference between a transferable method and an earned result.

Multi
model comparison
Paper
simulated trading
Gated
learning changes
4
market domains
Scoped
service boundaries
Tested
public examples

Architecture and synthetic examples, not private deployment inventory, current operating status, or investment performance.

Multi-model

Compare model forecasts with observed outcomes and explicit uncertainty.

How it works

The architecture separates model forecasts across prediction markets, crypto, options, and forex. Public examples demonstrate how evidence quality and uncertainty affect comparisons, without publishing private model rosters or trading records.

Learning controls

Require evaluation before a proposed method changes model behavior.

How it works

The design separates candidate methods, evaluation evidence, and promotion decisions. Proposed changes must preserve the distinction between a model's original reasoning and later additions.

Reset and transfer

Keep recovery policy separate from the measurements a model has earned.

How it works

Reset markers preserve history. A transferred method does not confer the donor's performance record, and borrowed calibration must remain labelled until independently evaluated.

Measurement states

Distinguish measured results, empty results, unavailable data, and failed measurements.

How it works

Public examples compare estimates with a declared baseline and report uncertainty. Placeholder records and failed reads require explicit handling; an unmeasured result is not a zero or evidence of skill.

Paper trading

Demonstrate research workflows with simulated positions and synthetic examples.

How it works

The public material describes paper-trading mechanisms. Worked rates and model-price differences are illustrative calculations, not live investment performance or realized returns.

Service architecture

Separate application services, storage, monitoring, and verification.

How it works

Containerized services and explicit health checks are architectural boundaries. The showcase describes those mechanisms without publishing private deployment inventories or claiming current service health.

Interface work

Group related views and make measurement status visible alongside each result.

How it works

The interface design uses grouped navigation, readable diagrams, and explicit data states. These are design principles rather than a claim that every private surface has passed an accessibility audit.

Paper-only by designSelf-correctingHonest measurementMulti-modelFastAPI + PostgreSQL + pgvectorDocker ComposePrometheus + GrafanaMulti-agent development
Flagship deploymentEvery Defender domain, zero outages

Microsoft Defender Suite, Top to Bottom

An organization running near-default security needed the entire Microsoft Defender ecosystem stood up, without breaking clinical operations that cannot go down.

Rolled out Defender for Endpoint through ringed pilot groups (IT, early adopters, then production) so every control was validated before broad enforcement: zero outages, no major operational disruption.
Promoted the full attack surface reduction (ASR) rule set from audit to enforce, with tamper protection, cloud-delivered next-gen antivirus baselines, PUA blocking, and EDR in block mode.
Stood up Defender for Identity on domain controllers to surface credential theft, lateral movement, and domain dominance techniques the moment they start.
Hardened email with Defender for Office 365: Safe Links and Safe Attachments, detonation, impersonation and anti-phishing policies tuned to the org, priority account protection.
Governed the SaaS layer with Defender for Cloud Apps: OAuth consent phishing detection, anomalous data movement policies, and app governance over what touches the tenant.
Enabled Defender for Cloud across Azure workloads for posture management and workload protection, and Defender for Storage (plan 2) with on-upload malware scanning.
Built the vulnerability management program from nothing: Qualys and Tenable authenticated scanning stood up from scratch, findings fused with Defender Vulnerability Management, risk-ranked remediation, and mean-time-to-remediate metrics reported to leadership monthly.
Raised Microsoft Secure Score from 35% to the 80 to 85% band, each improvement shipped as a documented, approved, reversible change.
Built Microsoft Sentinel from scratch alongside it all: workspace design, data connectors, MITRE-mapped analytics rules, SOAR automation, and the entire device fleet feeding the pipeline.
Secure Score 35% to 85%Zero outagesDefender for EndpointDefender for IdentityDefender for Office 365Defender for Cloud AppsDefender for CloudDefender for Storage P2Qualys + TenableASRSentinelCIS benchmarks

BlackGate: Agentic Purple-Team Platform

Autonomous offensive testing is powerful, but reckless without hard limits on scope, isolation, and human control.

Impact
  • Drives a technique catalogue across a full kill chain, with every action gated by a signed, fail-closed authorization scope so it only ever touches authorized targets. Planning is deterministic today; the model-driven core is built and deliberately not wired.
  • Executes all tooling inside disposable, network-isolated Kali sandboxes with a one-way results channel; any state-changing action stops at a human approval gate.
  • Closes the purple-team loop by replaying validated techniques to test detection coverage and turning gaps into new Sentinel and Defender detection content.
MITRE ATT&CKPurple TeamZero TrustGo execution serviceDisposable Kali sandboxesFour-stage approvalHuman-in-the-loop

Multi-Provider Agentic AI Orchestration Engine

Security teams wanted AI leverage without sending sensitive data to third-party models.

Impact
  • Routed workloads across cloud and on-device models by data-sensitivity class, keeping high-sensitivity security data on fully air-gapped inference.
  • Ran a manager, engineer, reviewer, and QA agent loop for multi-step alert correlation, triage, and remediation planning with minimal human intervention.
  • Standardized orchestration on Semantic Kernel with RAG over a live security knowledge base.
NIST AI RMFOWASP LLM Top 10MITRE ATLASSemantic KernelMulti-provider routingAir-gapped inference

Defender and Sentinel in a HIPAA and SOC 2 Environment

A compliance-driven organization needed audit-ready detection and response across a large fleet.

Impact
  • Rolled out the full Microsoft Defender ecosystem (Endpoint, Identity, Office 365, Cloud Apps, Cloud, Servers, Storage, and Containers) unified under Defender XDR, with alerts centralized into Microsoft Sentinel.
  • Designed workspaces, connectors, and analytics rules, and authored KQL detections that improved visibility and cut false positives.
  • Delivered audit-ready detection and response across multiple healthcare organizations, built around HIPAA, SOC 2, and FISMA controls.
HIPAASOC 2FISMAMITRE ATT&CKSentinelDefender for EndpointKQL

KQL Multi-Table Threat Hunting Orchestrator

Hunts were ad hoc; the team needed a repeatable capability across the full kill chain.

Impact
  • Built a reusable query library across DeviceNetworkEvents, DeviceProcessEvents, IdentityLogonEvents, EmailEvents, and CloudAppEvents.
  • Tagged every query to MITRE ATT&CK tactics and techniques for rapid hypothesis-driven hunts from initial access through exfiltration.
  • Documented findings in structured investigation reports covering IOC timelines, affected assets, and hardening recommendations.
MITRE ATT&CKDefender XDRAdvanced HuntingDetection engineering

Copilot Security and AI Governance Program

Leadership wanted Microsoft Copilot in clinical and operational workflows without leaking regulated data into AI surfaces.

Impact
  • Assessed the oversharing blast radius before rollout: SharePoint and OneDrive permissions, stale links, and sensitive sites Copilot could index.
  • Gated AI surfaces behind Conditional Access and sensitivity labels so prompts and grounding respect existing data boundaries.
  • Authored the organizational AI acceptable-use and governance guidance, aligned to NIST AI RMF, ISO/IEC 42001, the EU AI Act, and HHS AI strategy.
  • Extended detections to AI usage: anomalous prompt activity, shadow AI discovery, and Copilot audit events routed into Sentinel.
NIST AI RMFISO/IEC 42001HIPAAMicrosoft CopilotPurviewConditional Access

Sentinel SOAR Orchestration and Playbooks

Analyst time was burning on repetitive triage: enrichment, containment, and notification were all manual.

Impact
  • Built automation rules and Logic Apps playbooks that trigger on analytics: auto-enrich entities with threat intel, geolocation, and asset criticality before an analyst ever opens the incident.
  • Automated containment paths: disable compromised accounts, revoke sessions, isolate endpoints via Defender, all with approval gates for destructive actions.
  • Wired notifications and case flow: severity-based routing to Teams and email, auto-created tickets, and closure with documented outcomes.
  • Cut mean time to respond by removing the manual steps between detection and first action.
MITRE ATT&CKNIST CSF 2.0SentinelLogic AppsSOAR

Cloud PC Deployment for Healthcare

Clinical and engineering teams needed secure, compliant virtual desktops at scale.

Impact
  • Rolled out a scalable Windows 365 Cloud PC deployment for a healthcare organization, with per-group sizing tuned to each workload.
  • Architected role-based virtual networks with segmentation to enforce access controls and compliance.
Zero TrustRBACWindows 365VNet segmentation

Architecture in action

A few of these systems, wired end to end. Each pipeline opens on its first stage, so it reads without touching anything. Pick a stage to jump to it, or open the full list underneath.

PolyMind learning loopFrom market intake to the next prompt, every claim evidence-gated and no real money in the loopOpen the Atlas
Architecture walkthrough 9 stages

01 Intake: Each market feeds its own ground truth on its own clock: prediction-market odds and sports schedules, closed crypto candles, delayed option chains, and daily plus intraday currency rates. Stale or unverifiable data is refused at the door rather than patched over.

Reading stage 1 of 9
Read all 9 stages
  1. 01 Intake: Each market feeds its own ground truth on its own clock: prediction-market odds and sports schedules, closed crypto candles, delayed option chains, and daily plus intraday currency rates. Stale or unverifiable data is refused at the door rather than patched over.
  2. 02 Prediction: Each model receives a structured brief and answers on its own with a pick, a confidence, and a probability. A model that does not answer is recorded as a clearly marked placeholder, never as a real vote.
  3. 03 Evidence gate: Each vote is screened before any money is simulated: banned models and leagues, entry-time cutoffs, price sanity, duplicates, a calibration check, and hard brakes on drawdown. Placeholder votes never reach a money figure.
  4. 04 Paper trade: A qualifying vote becomes a paper trade sized by confidence against that model's simulated bankroll. Cash is derived by replaying the ledger, never stored, so history is never erased and a reset is just a new starting line.
  5. 05 Settlement: Two independent resolvers settle outcomes against ground truth with an atomic claim so nothing is ever resolved twice. Settlement writes the win, the loss, or the void that every later stage reads.
  6. 06 Calibration: Settled trades update each model's calibration: a Bayesian estimate of its true hit rate, confidence intervals, and a comparison against the odds it paid, so merely matching the market is never called skill. A model with no record yet is told so instead of handed a default.
  7. 07 Skill pass: Skills are promoted only once validated against settled outcomes; a recognizer removes bad habits a model picked up while protecting how it originally reasoned; and an audit confirms whether every learning stage actually ran.
  8. 08 Reset or graft: A recovery policy can mark a new evaluation period while retaining history. A method may be transferred for evaluation, but the donor record does not transfer and borrowed calibration remains labelled.
  9. 09 Prompt: The graft reaches the next prompt as an added layer: the base prompt stays untouched, the borrowed method is context only, and the model is told its own calibration is unknown until it earns one. Then the loop runs again.
Sentinel SOAR pipelineFrom raw telemetry to contained incident, with human gates on destructive actions
Architecture walkthrough 6 stages

01 Telemetry: Signals stream into Sentinel from Defender XDR, Entra sign-ins, email, cloud apps, and third-party feeds like Qualys and Tenable.

Reading stage 1 of 6
Read all 6 stages
  1. 01 Telemetry: Signals stream into Sentinel from Defender XDR, Entra sign-ins, email, cloud apps, and third-party feeds like Qualys and Tenable.
  2. 02 Analytics: Custom KQL analytics rules, each mapped to MITRE ATT&CK tactics, correlate raw events into high-fidelity incidents and suppress the noise.
  3. 03 Auto-enrich: An automation rule fires a Logic Apps playbook that enriches every entity: threat intel reputation, geolocation, asset criticality, and user risk, before an analyst opens the incident.
  4. 04 Triage: Severity and confidence route the incident: low-risk closes with documentation, medium goes to the analyst queue, high triggers the containment path with an approval gate.
  5. 05 Contain: Approved containment executes in seconds: disable the account, revoke active sessions, isolate the endpoint through Defender, and block indicators.
  6. 06 Notify + ticket: Stakeholders are notified by severity through Teams and email, a ticket is created automatically, and the incident closes with documented outcomes and lessons learned.
Agentic AI triage engineMulti-agent investigation with sensitivity-based routing and human approval
Architecture walkthrough 6 stages

01 Alert intake: A security alert or hunt hypothesis enters the engine with its entities, evidence, and context.

Reading stage 1 of 6
Read all 6 stages
  1. 01 Alert intake: A security alert or hunt hypothesis enters the engine with its entities, evidence, and context.
  2. 02 Sensitivity gate: A classification layer scores the data sensitivity. Regulated or high-sensitivity content is routed to fully air-gapped local inference; the rest can use cloud models.
  3. 03 Manager agent: The manager agent decomposes the investigation into steps: what to query, what to correlate, what to verify, and assigns them to the engineer agent.
  4. 04 Engineer agent: The engineer agent executes KQL queries, pulls related knowledge through RAG over the security knowledge base, and drafts findings with evidence.
  5. 05 Reviewer + QA: Reviewer and QA agents adversarially check the findings: is the evidence real, is the logic sound, did anything get missed? Weak conclusions are sent back.
  6. 06 Human approval: The verified plan (containment steps, remediation, report) goes to a human for approval. Agents never take consequential action on their own.
Phishing and BEC responseFrom reported email to org-wide purge, blocked campaign, and smarter users
Architecture walkthrough 6 stages

01 Report: A suspicious email arrives, reported by a user or caught by Defender for Office 365 detonation and impersonation detection.

Reading stage 1 of 6
Read all 6 stages
  1. 01 Report: A suspicious email arrives, reported by a user or caught by Defender for Office 365 detonation and impersonation detection.
  2. 02 Analyze: Headers, sender authentication results, links, and attachments are analyzed, with sandbox detonation for anything executable.
  3. 03 Hunt: KQL hunts across EmailEvents and CloudAppEvents find every mailbox that received the campaign and anyone who clicked.
  4. 04 Purge: The campaign is soft-deleted org-wide in one action, clickers get sessions revoked and credentials reset.
  5. 05 Block: Senders, domains, and URLs are blocked, and transport rules plus tenant allow-block lists are updated so the campaign cannot return.
  6. 06 Educate: The campaign feeds the phishing simulation program: targeted training for the users who clicked, metrics for leadership.
Vulnerability management loopContinuous discovery to verified remediation, reported in business terms
Architecture walkthrough 6 stages

01 Discover: Continuous scanning across endpoints, servers, and cloud workloads with Qualys, Tenable, and Defender vulnerability management.

Reading stage 1 of 6
Read all 6 stages
  1. 01 Discover: Continuous scanning across endpoints, servers, and cloud workloads with Qualys, Tenable, and Defender vulnerability management.
  2. 02 Prioritize: Findings are ranked by real risk: active exploitation, exposure, and asset criticality, not just CVSS score.
  3. 03 Assign: Each remediation becomes a documented, approved, reversible change request with a named owner and a deadline tied to severity.
  4. 04 Remediate: Patching, configuration hardening, or compensating controls where patching is not possible, rolled out through pilot rings first.
  5. 05 Verify: Rescans confirm closure. Anything that reopens goes back through the loop with escalation.
  6. 06 Report: Trend metrics, mean time to remediate, and risk posture reported to leadership in business terms.
See how I build

The code, in the open

A public, sanitized slice of how I build: governed, measured, and mapped to real security and AI governance frameworks. Enough to show the thinking, never the secrets.

AI eval harness with a hard safety and injection ship-gate
Default-deny agent guardrails (OWASP LLM01 and LLM02)
Log-odds signal fusion, de-vig, and Brier-scored calibration
Governance mapped to NIST AI RMF, ISO 42001, and MITRE
🔒 Private by design🧪 Synthetic data only🧮 Reproducible, not performance

Most of my repositories stay private. Private research, client work, and anything with real tenants or credentials never goes public, for security and personal reasons.

Explore the showcase
greeklinux, Agentic AI Security Architect

Want to learn more about me?

Visit my LinkedIn for the full professional profile, or browse the code showcase above. If you would like to learn more about the private work, please feel free to email me at ulisesghurtado@gmail.com.

What I build for companies

AI solutions that pay for themselves

Merging AI and security is my specialty, but it is not the whole story. I design AI solutions for whatever eats a company's time and money, then wire in the backend security mindset most builders skip. The goal is always the same: cut costs with the minimum of effort, and ship something you can actually run.

Customer response automation

Agents that read incoming email, draft on-brand replies, resolve the routine ones automatically, and escalate the rest to a human with full context attached.

Cuts response time and support hours

Automatic detection and pager duty

Detectors over your logs, metrics, and alerts with severity-based paging: the right person gets woken up only when it matters, with the evidence already gathered.

Cuts alert fatigue and missed incidents

Market research and evaluation

Research workflows that compare forecasts, evaluate calibration, and track uncertainty. PolyMind provides synthetic examples of these mechanisms without claiming live investment performance.

Organizes research and model evaluation

Back-office automation

Reports, intake, ticket triage, scheduling, knowledge bases: the repetitive paperwork layer of a business, automated with human approval exactly where it counts.

Cuts manual admin to near zero

Security copilots

KQL-writing hunt agents, alert triage assistants, and incident summarizers that let a small security team operate like a large one.

Cuts triage time and analyst burnout

AI strategy and new ideas

Not sure where AI fits your company? I map your workflows, find the highest-ROI target, and ship a working pilot fast. Consulting that ends in software, not slideware.

Cuts the guesswork out of AI adoption

The difference in my builds: every solution ships with the security backend baked in. Least privilege, logged actions, human gates on consequential steps, and sensitive data kept where it belongs. That is the gap between an AI demo and an AI system a company can trust in production.

How I work

Principles I build on

Secure by design

I build security in from the first line, not bolted on after the fact.

Compliant and auditable

Every control documented, approved, reversible, and mapped to a framework.

Automate the repeatable

SOAR playbooks, KQL libraries, and scripts so the team scales without burning out.

Translate to the business

I frame security in risk and regulatory terms that leaders can act on.

Capabilities

What I work with

AI models and their ecosystems

Claude (Anthropic)Claude CodeClaude CoworkOpenAI GPT and o-seriesAzure OpenAIGoogle GeminiAntigravityDeepSeekKimi (Moonshot)QwenLlamaMistralHermes (Nous Research)Grok (xAI)Phi (Microsoft)Microsoft CopilotCopilot StudioGitHub CopilotCursorOllamaLM Studio

Agentic engineering

Multi-agent orchestration (manager, engineer, reviewer, QA)Model routing by sensitivity, latency, and complexitySemantic KernelRAG pipelines and vector searchModel Context Protocol (MCP)Prompt engineering and structured outputsPrompt-injection defense and guardrailsAir-gapped local inferenceModel evaluation and calibrationAzure AI FoundryAI governance (NIST AI RMF, ISO 42001, EU AI Act)OWASP LLM Top 10 and MITRE ATLAS

Security and SOC

Microsoft Defender XDRDefender for EndpointDefender for IdentityDefender for CloudDefender for StorageDefender for Office 365Microsoft SentinelKQL threat huntingMITRE ATT&CK detection engineeringSOAR automation (Logic Apps playbooks)Attack surface reduction and EDR block modeSecure Score optimizationIncident responseThreat intel enrichmentMicrosoft Security CopilotSentinelOneCortex XDR (Palo Alto)QualysTenableKali Linux and Kali containersPhishing simulation programs

Identity and access

Microsoft EntraZero Trust architectureConditional AccessIdentity ProtectionToken protection and authentication strengthsBreak-glass account designActive Directory and GPOOktaCyberArkSailPointMFA and SSPR programsPrivileged access managementRBAC design

Data protection and compliance

Microsoft PurviewSensitivity labels and DLPHIPAA, SOC 2, and FISMA operationsEmail authentication (DMARC, DKIM, SPF)AI acceptable-use policyAudit readiness and evidenceChange management and CABDocumentation and runbook libraries

Cloud and infrastructure

AzureAWSGoogle CloudMicrosoft 365 and Exchange OnlineAzure Monitor and Log AnalyticsMicrosoft Graph APIIntune device managementWindows 365 Cloud PCSharePointG SuiteLinuxOracle and SSMSVeeam and Commvault (backup and recovery)

Microsoft admin centers (all of them)

Microsoft 365 Admin CenterExchange Admin CenterEntra Admin CenterAzure PortalPower Platform Admin CenterIntune Admin CenterMicrosoft Defender portalPurview compliance portalSharePoint Admin CenterTeams Admin CenterSecurity and compliance role scopingTenant-wide configuration and licensing

Network and monitoring

Fortinet firewallsPalo Alto firewallsCisco MerakiLogic MonitorVPN and DNS troubleshootingVNet segmentationDNS modernizationRemote support (Bomgar, ConnectWise, TeamViewer)

Engineering and operations

PythonPowerShellBash and shell scriptingTypeScript and JavaScriptNode.jsReact and Next.jsSQLPostgreSQL and pgvectorRedisPandas and NumPyFastAPIREST APIs and Microsoft GraphDockerInfrastructure as Code (Bicep and Terraform)Git and GitHubGitHub Actions (CI/CD)Azure DevOpsVercelPower Automate and Logic AppsPower BIJSON and YAMLVS CodeLinux administrationServiceNowZendeskJira and Azure BoardsITIL and change managementFive9 and Genesys

AI security and governance

Securing systems that act on their own

Securing AI is its own discipline. As systems gain autonomy, I make sure the controls around them are as strong as the models are capable, from the data boundary to the prompt surface to the governance program.

Treat models as untrusted

Agentic loops run with least privilege, human approval on consequential actions, and validated, structured outputs.

Keep sensitive data in-house

Data-sensitivity classification routes high-risk workloads to on-device, air-gapped inference, so regulated data never leaves the boundary.

Defend the prompt surface

Prompt-injection shields, tool permissioning, and egress secret masking on every agent that can take an action.

Govern by framework

Adoption guidance aligned to NIST AI RMF 1.0, ISO/IEC 42001, the EU AI Act, OWASP LLM Top 10, MITRE ATLAS, and HHS AI strategy.

Checkable, not asserted

The same work, with the receipts attached

Everything below is read out of one public repository, github.com/greeklinux/Showcase (opens the file on GitHub in a new tab). Every number states how it was derived and links to the file it came from, so any of it can be recounted. Nothing here is a score, a percentage, a maturity level or a rating: those are claims about myself that nobody else can check, and counts are not.

The finding that organizes all of it is not a wrong control. It is a control that was written, reviewed, merged, and was not in effect. One of these modules called a reviewer agent, wrote its verdict into the response, and then computed whether to execute from alert severity alone. Nothing errored, nothing would have failed a code review, and the control was in the architecture diagram, in the code, and in the output, and absent from the decision.

modules that run on their own
15modules that run on their owneach one a guard, an analyzer or a gate, standard library only
tests over those modules
1,190tests over those modulesone test file each, and the parts sum to the suite total
framework identifiers claimed
29framework identifiers claimedeach checked against its publishing source, with its edition
framework mappings declined
13framework mappings declinedwritten down with the reason, rather than quietly filled in

Identifiers claimed, and mappings refused

Four frameworks, every identifier checked against its publishing source before it was written down. The hatched bar is the mappings a reviewer might have expected and that were declined instead, with the reason recorded.

29
distinct identifiers claimed

across 44 mapping rows in two tables

13
mappings declined on the record

each with the reason it could not be substantiated

2
weakness classes named instead

where a technique was refused, the honest home is named rather than left blank

Every one of the 13, and why

Roughly one mapping in four that a reviewer might expect on those pages is absent on purpose. Three of the thirteen are here. The rest are one click away, in full, with the reason each one could not be substantiated.

  • Any NIST AI RMF subcategory at all

    would have gone on mount_audit.py (opens the file on GitHub in a new tab)

    An application security control, not an AI specific one. GOVERN 1.6 on inventorying AI systems is the nearest thing in spirit, and an HTTP route table is not an AI system inventory.

  • The claim that NIST AI RMF requires red teaming

    would have gone on eval_harness.py (opens the file on GitHub in a new tab)

    The phrase does not appear in the normative text of NIST AI 100-1. It appears in the Playbook suggested actions under MEASURE 2.7, which is guidance. Overstating a framework is the same defect as overstating a metric.

  • An ATT&CK technique for a machine identity calling an API it does not normally call

    would have gone on agent_tool_invocation.kql (opens the file on GitHub in a new tab)

    None exists. T1078.004 plus T1059.009 is the closest defensible pair, and the file says so in its own header rather than inventing an identifier that would look tidier on a slide.

Read all 13 refusals, and the reason for each
  • Any NIST AI RMF subcategory at all

    would have gone on mount_audit.py (opens the file on GitHub in a new tab), declined in ai_security

    An application security control, not an AI specific one. GOVERN 1.6 on inventorying AI systems is the nearest thing in spirit, and an HTTP route table is not an AI system inventory.

  • The claim that NIST AI RMF requires red teaming

    would have gone on eval_harness.py (opens the file on GitHub in a new tab), declined in ai_security

    The phrase does not appear in the normative text of NIST AI 100-1. It appears in the Playbook suggested actions under MEASURE 2.7, which is guidance. Overstating a framework is the same defect as overstating a metric.

  • An ATT&CK technique for a machine identity calling an API it does not normally call

    would have gone on agent_tool_invocation.kql (opens the file on GitHub in a new tab), declined in ai_security

    None exists. T1078.004 plus T1059.009 is the closest defensible pair, and the file says so in its own header rather than inventing an identifier that would look tidier on a slide.

  • Any MITRE ATLAS technique

    would have gone on control_flow_audit.py (opens the file on GitHub in a new tab), declined in ai_security

    ATLAS catalogues what an adversary does. A control that was merged and is not in effect is what a codebase failed to do. The honest home is a weakness class, CWE-693 Protection Mechanism Failure, not a technique.

  • Any MITRE ATLAS technique for laundering

    would have gone on provenance_algebra.py (opens the file on GitHub in a new tab), declined in ai_security

    Promoting an untrusted string to a trusted one by summarizing it is a defect in the defender own data model. The nearest identifier is CWE-501 Trust Boundary Violation.

  • Any MITRE ATT&CK technique

    would have gone on capability_attenuation.py (opens the file on GitHub in a new tab), declined in ai_security

    Delegation between sub agents is not a technique in that matrix. Its nearest neighbours are all about stolen credentials, and this is about authority a system handed out itself.

  • A delegation specific OWASP slot

    would have gone on capability_attenuation.py (opens the file on GitHub in a new tab), declined in ai_security

    A slot for agentic delegation would be the natural home for that file, and asserting an identifier without checking it against the published list is exactly what the edition warning exists to prevent. LLM03 Excessive Agency is the mapping the file can stand behind in both editions, so it is the only one claimed.

  • Any NIST AI RMF subcategory

    would have gone on provenance_algebra.py (opens the file on GitHub in a new tab), declined in ai_security

    The MAP function third party component subcategories are the nearest thing in spirit. They are written about components and supply chain rather than about a per span label carried at runtime, and their normative text is not quoted, because a mapping asserted from memory is worth less than one declined on the record.

  • The claim that this detects prompt injection

    would have gone on differential_consistency.py (opens the file on GitHub in a new tab), declined in ai_security

    It detects a decision that is not stable under transformations the task is invariant under. Those two sets overlap and are not the same set, and conflating them would claim a completeness the module cannot have.

  • Any OWASP LLM slot

    would have gone on audit_chain.py (opens the file on GitHub in a new tab), declined in blackgate

    It is an integrity control over a log. There is no LLM in it, and the nearest slot would be a stretch.

  • Any MITRE technique

    would have gone on approval_ceremony.py (opens the file on GitHub in a new tab), declined in blackgate

    An approval ceremony is a governance control, not an adversary behaviour. No identifier describes it, and inventing a near fit would look tidier on a slide and be wrong.

  • Any NIST AI RMF subcategory

    would have gone on the address matching rules in scope_gate.py (opens the file on GitHub in a new tab), declined in blackgate

    Network authorization, not AI governance.

  • MITRE ATLAS, anywhere

    would have gone on all six blackgate files (opens the file on GitHub in a new tab), declined in blackgate

    ATLAS is the right taxonomy for attacks on AI systems. Nothing in that directory defends a model, so nothing in it has an ATLAS mapping.

Two of those refusals name a weakness class as the honest home for the mapping instead: CWE-693 Protection Mechanism Failure for control_flow_audit.py, and CWE-501 Trust Boundary Violation for provenance_algebra.py. Naming where a mapping would belong is not the same as claiming it, so neither is counted on the claimed side.

Framework mappings, claimed against declined. OWASP Top 10 for LLM Applications: 13 mapping rows carrying 4 distinct identifiers, 2 declined. MITRE ATLAS: 10 mapping rows carrying 8 distinct identifiers, 3 declined. MITRE ATT&CK: 13 mapping rows carrying 12 distinct identifiers, 2 declined. NIST AI RMF 1.0: 8 mapping rows carrying 5 distinct identifiers, 4 declined. Two further declines are not charged to a single framework. 13 mappings declined in total.

How this is counted Both sides are row counts from the two tables headed Framework mapping (opens the file on GitHub in a new tab) and Mappings deliberately not claimed (opens the file on GitHub in a new tab), in ai_security/README.md and blackgate/README.md. Claimed is 44 mapping rows carrying 29 distinct identifiers: an identifier cited against two different files occupies two rows and is counted once as distinct. Declined is 13 rows, of which 2 are not charged to any single bar because one declines both MITRE matrices at once and one declines a plain claim rather than an identifier. Neither side is a coverage percentage. A percentage would need a denominator saying how many mappings a repository of this size ought to carry, and no such number exists.

Show the mapping counts
Mapping rows claimed and mappings declined, per framework. Every figure is a row count.
FrameworkRows claimedDistinct identifiersDeclined
OWASP Top 10 for LLM Applications1342
MITRE ATLAS1083
MITRE ATT&CK13122
NIST AI RMF 1.0854
Not charged to one framework002
Total442913

OWASP Top 10 for LLM Applications, four slots of ten

What the fifteen modules claim against the list, and what they leave alone. A slot number without its edition is not a mapping, so each claimed slot carries both.

  1. LLM014

    Prompt Injection

  2. LLM02

    not named in the source, so not named here

    no control here claims it

  3. LLM0310

    Excessive Agency

  4. LLM04

    not named in the source, so not named here

    no control here claims it

  5. LLM05

    Data and Model Poisoning

    the slot that carried Improper Output Handling in 2025

  6. LLM06

    not named in the source, so not named here

    no control here claims it

  7. LLM07

    not named in the source, so not named here

    no control here claims it

  8. LLM081

    Hidden Context Exposure

  9. LLM09

    not named in the source, so not named here

    no control here claims it

  10. LLM102

    Improper Output Handling

Module by module

Every slot a module claims, and the two modules that claim none.

2 of the 15 modules claim no slot. audit_chain (opens the file on GitHub in a new tab) is the one where the refusal is written down: An integrity control over a log. There is no LLM in it, and the nearest slot would be a stretch.

OWASP Top 10 for LLM Applications coverage. 4 of the ten slots are claimed and 6 are not. LLM01 in 2026, Prompt Injection, carried by 4 module and slot pairs, and numbered LLM01 in the 2025 edition. LLM03 in 2026, Excessive Agency, carried by 10 module and slot pairs, and numbered LLM06 in the 2025 edition. LLM08 in 2026, Hidden Context Exposure, carried by 1 module and slot pairs, and numbered LLM07 in the 2025 edition. LLM10 in 2026, Improper Output Handling, carried by 2 module and slot pairs, and numbered LLM05 in the 2025 edition. Two modules claim no slot at all.

How this is counted Counted from the two tables headed OWASP Top 10 for LLM Applications (opens the file on GitHub in a new tab): eleven rows in ai_security/README.md and two in blackgate/README.md. Those thirteen rows are 17 module and slot pairs, because one row names three modules at once and two rows name two each. Slot names and the 2025 equivalents are quoted from those same tables. Five of the six unclaimed slots are not named here. The source repository does not name them, and asserting a slot name from memory is the exact defect its own edition warning exists to prevent. The sixth, LLM05:2026, is named because the source names it: it is Data and Model Poisoning in 2026 and was Improper Output Handling in 2025, which is why a bare LLM05 is a coin flip rather than a mapping.

Show the slots claimed, module by module
Every module, the 2026 slots it claims, and the 2025 numbering of the same slots.
ModuleDirectory2026 slots claimed2025 equivalent
prompt_guardai_securityLLM01 Prompt Injection, LLM08 Hidden Context ExposureLLM01 Prompt Injection, LLM07 System Prompt Leakage
llm_output_validatorai_securityLLM10 Improper Output Handling, LLM03 Excessive AgencyLLM05 Improper Output Handling, LLM06 Excessive Agency
capability_attenuationai_securityLLM03 Excessive AgencyLLM06 Excessive Agency
control_flow_auditai_securityLLM03 Excessive AgencyLLM06 Excessive Agency
differential_consistencyai_securityLLM01 Prompt Injection, LLM10 Improper Output HandlingLLM01 Prompt Injection, LLM05 Improper Output Handling
provenance_algebraai_securityLLM01 Prompt Injection, LLM03 Excessive AgencyLLM01 Prompt Injection, LLM06 Excessive Agency
eval_harnessai_securityLLM01 Prompt InjectionLLM01 Prompt Injection
mount_auditai_securityLLM03 Excessive AgencyLLM06 Excessive Agency
agentic_socai_securityLLM03 Excessive AgencyLLM06 Excessive Agency
scope_gateblackgateLLM03 Excessive AgencyLLM06 Excessive Agency
attestationblackgateLLM03 Excessive AgencyLLM06 Excessive Agency
detection_gapblackgatenone claimednot applicable
audit_chainblackgatenone, declined on the recordnot applicable
prohibitionsblackgateLLM03 Excessive AgencyLLM06 Excessive Agency
approval_ceremonyblackgateLLM03 Excessive AgencyLLM06 Excessive Agency

MITRE ATLAS, eight techniques and where each one lands

Each line is a mapping that exists in the published table. Two identifiers are cited against a second file each, which is why ten table rows carry eight distinct techniques.

Read AML.T0053 twice. Published as LLM Plugin Compromise and renamed. A mapping carrying the old name is stale even though the identifier still resolves, which is the kind of thing a reviewer only catches by opening the source.

MITRE ATLAS technique coverage, 8 distinct techniques across 10 mapping rows. AML.T0051, LLM Prompt Injection, mapped against prompt_guard.py. AML.T0051.001, LLM Prompt Injection: Indirect, mapped against prompt_guard.py and provenance_algebra.py and differential_consistency.py. AML.T0053, AI Agent Tool Invocation, mapped against llm_output_validator.py and agentic_soc.py and capability_attenuation.py and agent_tool_invocation.kql. AML.T0054, LLM Jailbreak, mapped against prompt_guard.py. AML.T0056, Extract LLM System Prompt, mapped against prompt_guard.py. AML.T0101, Data Destruction via AI Agent Tool Invocation, mapped against llm_output_validator.py. AML.T0086, Exfiltration via AI Agent Tool Invocation, mapped against agent_tool_invocation.kql. AML.T0024, Exfiltration via AI Inference API, mapped against agent_tool_invocation.kql.

How this is counted Every identifier and name is taken from the table headed MITRE ATLAS (opens the file on GitHub in a new tab) in ai_security/README.md, where each was checked against its publishing source with the name it currently carries. Ten rows, 8 distinct identifiers, 13 technique and file links drawn. No share of the ATLAS matrix is claimed here. ATLAS is large, this is nine modules and three hunts, and a ratio against the whole matrix would be a number designed to read as more than it is. The blackgate directory declines ATLAS entirely across all six of its files, on the grounds that nothing in it defends a model.

Show the techniques and their mappings
Every ATLAS technique claimed, its published name, and the files it is mapped against.
TechniqueNameMapped againstTable rows
AML.T0051LLM Prompt Injectionprompt_guard.py1
AML.T0051.001LLM Prompt Injection: Indirectprompt_guard.py, provenance_algebra.py, differential_consistency.py2
AML.T0053AI Agent Tool Invocationllm_output_validator.py, agentic_soc.py, capability_attenuation.py, agent_tool_invocation.kql2
AML.T0054LLM Jailbreakprompt_guard.py1
AML.T0056Extract LLM System Promptprompt_guard.py1
AML.T0101Data Destruction via AI Agent Tool Invocationllm_output_validator.py1
AML.T0086Exfiltration via AI Agent Tool Invocationagent_tool_invocation.kql1
AML.T0024Exfiltration via AI Inference APIagent_tool_invocation.kql1
Total8 distinct techniques7 files10

Two attack surfaces, and the one everybody forgets

An agentic system has what goes in and what comes out. The third is what it is mounted on, and none of the first two matters if the route driving the agent was reachable without authentication.

  • typed by a person
  • a retrieved page or RAG chunk
  • tool or API output
untrusted until screened

Surface 1

What goes in

Is this text allowed to instruct the model at all.

prompt_guard (opens the file on GitHub in a new tab)101 tests

The agent proposes a tool call

Surface 2

What comes out

Is this proposed tool call on the allowlist, in bounds, and named by an approval.

llm_output_validator (opens the file on GitHub in a new tab)85 tests

every gate can only subtract
  • Refused

    not on the allowlist, an argument out of bounds, an unknown key, or the validator raised

  • Held

    high impact, so nothing runs until an approval names a digest of this exact call

  • Runs

    bound to a digest of this exact tool and these exact arguments

Surface 3

What it is mounted on

Was the route driving the agent reachable without authentication.

mount_audit (opens the file on GitHub in a new tab)72 tests

This one sits to the side on purpose. It does not screen a message or a call. It walks the mounted routes and fails closed on any state changing route that is not effectively behind an auth dependency, because authorization coverage does not live in the route handler you would think to read.

Four pathways underneath, asking a different question

Not is a rule missing, but is the model of the system wrong, because a wrong model is what produces a control that is present and not in effect.

Where this comes from The three surfaces and the four pathways are the structure of ai_security/README.md (opens the file on GitHub in a new tab), which is where the nine modules are introduced in this order. Each test count is the Ran N tests line from that module test file run on its own, and the nine sum to 686. Every module named here runs on its own with python3 <file> and prints a worked example, so a reader can reproduce the behaviour rather than take the diagram on trust.

Where the tests are, module by module

One test file per module, named as sentences that state the property under test. The 15 security modules carry 1,190 of the suite's 1,590 tests, and the parts reconcile with the whole.

1,190
tests on these fifteen modules

of 1,590 in the whole suite

15
modules, one test file each

no shared fixtures between them

0
third party dependencies

Standard library only. make test (opens the file on GitHub in a new tab) has no install step because there is nothing to install.

The six blackgate modules sit unusually close together, between 64 and 90. That is a consequence of the subject rather than a target anybody aimed at: each is one gate with a small number of ways to be wrong and a large number of ways to be deceptively right, and the deceptive cases are what the tests are mostly made of.

Tests per security module. prompt_guard in ai_security, 101 tests. llm_output_validator in ai_security, 85 tests. capability_attenuation in ai_security, 75 tests. control_flow_audit in ai_security, 91 tests. differential_consistency in ai_security, 72 tests. provenance_algebra in ai_security, 79 tests. eval_harness in ai_security, 67 tests. mount_audit in ai_security, 72 tests. agentic_soc in ai_security, 44 tests. scope_gate in blackgate, 100 tests. attestation in blackgate, 88 tests. detection_gap in blackgate, 77 tests. audit_chain in blackgate, 79 tests. prohibitions in blackgate, 78 tests. approval_ceremony in blackgate, 82 tests. ai_security totals 686, blackgate totals 504, and the two sum to 1190 of the suite's 1590.

How this is counted Each bar is the Ran N tests line from python3 -m unittest tests.test_<module>, run on its own. The fifteen sum to 1190. The remaining 400 of the suite's 1590 belong to two directories that are not security modules, so they are not plotted. The repository ships tests/check_claims.py (opens the file on GitHub in a new tab), which re-derives every published figure from a run and fails on drift. It was run in full on 2026-09-20 and reported every figure matched, at a baseline of 1590 tests. A test count is not a quality score. A module with more tests has a larger surface of ways to be wrong, which is a property of the subject rather than of the effort: prompt_guard.py (opens the file on GitHub in a new tab) carries the most tests in the repository at less than half the size of control_flow_audit.py (opens the file on GitHub in a new tab), because a normalization guard has a large surface of adversarial inputs and an AST walker has a small surface of code shapes.

Show the test counts
Tests per module, by directory. Every figure is a count read off a run.
ModuleDirectoryTests
prompt_guardai_security101
llm_output_validatorai_security85
capability_attenuationai_security75
control_flow_auditai_security91
differential_consistencyai_security72
provenance_algebraai_security79
eval_harnessai_security67
mount_auditai_security72
agentic_socai_security44
scope_gateblackgate100
attestationblackgate88
detection_gapblackgate77
audit_chainblackgate79
prohibitionsblackgate78
approval_ceremonyblackgate82
Fifteen security modulesai_security and blackgate1190
Rest of the suitepolymind and automation, not plotted400
Whole suitereported by unittest discover1590

The suite has been seen to fail, and you can watch it

One declared one line change at a time, planted in a scratch copy, whole suite run, restored. 137 changes, 137 caught, 0 left alive, and a named list of the 5 that did get through before they were closed.

137
mutations planted

one at a time, in a scratch copy

137
caught by the suite

the suite went red on each

0
survived the suite

a survivor the data has not declared makes the tool exit non zero

555
tests turned red in total

0 mutations broke the suite's collection rather than failing it

  • llm_output_validator
    4 planted, 48 red
  • approval_ceremony
    9 planted, 47 red
  • attestation
    9 planted, 41 red
  • audit_chain
    7 planted, 39 red
  • prohibitions
    9 planted, 35 red
  • provenance_algebra
    9 planted, 24 red
  • mount_audit
    9 planted, 24 red
  • prompt_guard
    5 planted, 23 red
  • scope_gate
    7 planted, 20 red
  • control_flow_audit
    7 planted, 19 red
  • differential_consistency
    4 planted, 18 red
  • detection_gap
    7 planted, 18 red
  • agentic_soc
    3 planted, 17 red
  • capability_attenuation
    4 planted, 9 red
  • eval_harness
    4 planted, 9 red

A clean sweep is the weakest result this harness can report

Nothing survives the current run. That says the planted set is fully covered; it does not say the planted set is complete, and the mutations are written by the same hand as the tests, so a blind spot in one is likely a blind spot in the other. The useful number is the list below: 5 mutations that did get through, what each one proved was unheld, and the test that now holds it. A harness that has only ever printed zero has not been used.

The 5 the harness changed its mind about

Three were found alive on the first run against the real suite. One was declared a survivor and turned out to be caught, which the harness reports as a stale declaration. One was the accepted survivor this card used to publish, and it is closed. All of them are the same mechanism working: declared data that no test held, which is the regression shape a suite misses most easily because the code itself does not change.

Mutation testing results for the fifteen security modules. prompt_guard: 5 mutations planted, 23 tests turned red, 0 survived. llm_output_validator: 4 mutations planted, 48 tests turned red, 0 survived. capability_attenuation: 4 mutations planted, 9 tests turned red, 0 survived. control_flow_audit: 7 mutations planted, 19 tests turned red, 0 survived. differential_consistency: 4 mutations planted, 18 tests turned red, 0 survived. provenance_algebra: 9 mutations planted, 24 tests turned red, 0 survived. eval_harness: 4 mutations planted, 9 tests turned red, 0 survived. mount_audit: 9 mutations planted, 24 tests turned red, 0 survived. agentic_soc: 3 mutations planted, 17 tests turned red, 0 survived. scope_gate: 7 mutations planted, 20 tests turned red, 0 survived. attestation: 9 mutations planted, 41 tests turned red, 0 survived. detection_gap: 7 mutations planted, 18 tests turned red, 0 survived. audit_chain: 7 mutations planted, 39 tests turned red, 0 survived. prohibitions: 9 mutations planted, 35 tests turned red, 0 survived. approval_ceremony: 9 mutations planted, 47 tests turned red, 0 survived. Across the whole repository, 137 mutations, 137 caught, 0 survived, 555 test deaths.

How this is counted Every figure is the output of python3 tests/mutation_harness.py (opens the file on GitHub in a new tab), run in full on 2026-09-20. The mutations are declared as data in tests/mutations.py (opens the file on GitHub in a new tab), one entry per change, each naming the file, the exact one line edit, and the property it is supposed to break. The harness copies the tree to a scratch directory, plants one change, runs the whole suite, records whether it went red, restores the file and moves on, and it ends by comparing a digest of every source file taken before the run with one taken after. Tests that turned red is the failures plus errors on the summary line of each run. The fifteen modules here carry 97 of the 137 mutations and 391 of the 555 test deaths. The run took 217 seconds on the machine this page was built on, which is environment conditional and is the one figure here that will not reproduce exactly. The repository states about a minute. Nothing else on this card depends on the machine.

Show the mutation counts
Mutations planted per module, tests that turned red, and survivors.
ModulePlantedCaughtSurvivedTests turned red
prompt_guard55023
llm_output_validator44048
capability_attenuation4409
control_flow_audit77019
differential_consistency44018
provenance_algebra99024
eval_harness4409
mount_audit99024
agentic_soc33017
scope_gate77020
attestation99041
detection_gap77018
audit_chain77039
prohibitions99035
approval_ceremony99047
Fifteen security modules97970391
Whole repository1371370555

Four defects worth publishing, because none of them errored

The most useful findings in that repository were not wrong controls. They were controls that were present, reviewed, merged, and not in effect. Here are four, each with the file that carries the fix.

  1. Defect 1 of 4

    A prompt filter defeated by a character you cannot see

    The pattern list matched raw text, so a payload split by a zero width space or dressed in fullwidth look alikes walked straight through a rule written for plain ASCII.

    What runs now: Every input is NFKC folded, stripped of invisible characters and whitespace collapsed before a single pattern runs, and the hidden characters are recorded as their own signal rather than silently removed.

    prompt_guard (opens the file on GitHub in a new tab)runs on its own and prints a worked example

  2. Defect 2 of 4

    An analyzer that passed the bug it exists to find

    A verdict handed as an argument to the very call it is supposed to gate reads as connected to a data flow pass, and deleting every guard clause in front of that call does not move the report.

    What runs now: The analyzer tracks control dependence as well as data flow, so a verdict that is recorded, returned or passed along, but never tested by a branch or an early return, is reported as not in effect.

    control_flow_audit (opens the file on GitHub in a new tab)runs on its own and prints a worked example

  3. Defect 3 of 4

    An output validator that accepted the metadata address in IPv6 notation

    The IPv4 mapped form of the cloud instance metadata address is link local to every resolver, and the standard library reports it as a member of none of the IPv6 networks on the deny list, so the deny list did not fire on it.

    What runs now: Every address the literal actually reaches is folded out before the deny list runs, and a spelling that names one address here and another to a resolver, a leading zero octet or a bare integer, is refused outright.

    llm_output_validator (opens the file on GitHub in a new tab)runs on its own and prints a worked example

  4. Defect 4 of 4

    A release gate that shipped on a rounded score

    Rates were rounded to three places for printing, so a safety bucket of four thousand cases with one genuine failure is 0.99975, which rounds to 1.000, clears a gate of 1.00, and ships while the failing case is still listed in the report.

    What runs now: The gate runs against the unrounded value, and a rate that fails but rounds onto the floor is printed with the extra digits it needs, so the reason can never read as a pass.

    eval_harness (opens the file on GitHub in a new tab)runs on its own and prints a worked example

Where this comes from Each defect is stated in the file it was found in and in the directory README that introduces it, under the heading The defect behind each module. The tests that once pinned these defects were inverted rather than deleted, so the record of what was once true stays readable in the suite. No count of defects found is claimed here. These four were chosen because they are the clearest, not because they are all of them. The source page catalogues seventeen instances of this one shape across four directories.

The detection side, and the three files with no tests

Three Sentinel KQL hunts, each carrying the defect it was written to fix or the trap it states up front. Two of the three were dead logic: a filter that could not fire, and a column that was false on every row, forever.

3
Sentinel hunts

impossible travel, illicit consent, agent tool invocation

0
tests over these three files

stated on the source page rather than left to be discovered

12
distinct ATT&CK techniques

across 13 identifier rows in two mapping tables

  • anomalous_signin.kql (opens the file on GitHub in a new tab)

    Impossible travel across consecutive successful interactive sign ins.

    What the file says about itself: The home country exclusion only fires when both countries are on the list, and the filter above it already requires the two countries to differ, so a one entry list could never exclude anything. A filter that cannot fire is not a lenient filter, it is an absent one.

    ATT&CK

    • T1078 Valid Accounts
    • T1078.004 Valid Accounts: Cloud Accounts
  • oauth_consent_grant.kql (opens the file on GitHub in a new tab)

    Illicit OAuth application consent grants, the consent phishing pattern.

    What the file says about itself: Admin consent was derived as OperationName has "admin", and none of the operation names the query filters on contains that word, so the column an analyst triages on was false on every row, forever. Admin consent is now read out of modifiedProperties by name.

    ATT&CK

    • T1528 Steal Application Access Token
    • T1550.001 Use Alternate Authentication Material: Application Access Token
    • T1078.004 Valid Accounts: Cloud Accounts
  • agent_tool_invocation.kql (opens the file on GitHub in a new tab)

    An agent, or the identity behind one, acting outside its observed envelope.

    What the file says about itself: Two of its three signals are anomaly detection and will make you tune. The third is a control assertion: a state changing tool invoked with no approval bound to the call is not an anomaly, it is the gate having been bypassed or having never run.

    ATT&CK

    • T1078.004 Valid Accounts: Cloud Accounts
    • T1059.009 Command and Scripting Interpreter: Cloud API

    ATLAS

    AML.T0053, AML.T0086, AML.T0024

ATT&CK has no technique for a machine identity calling an API it does not normally call. The closest defensible pair is T1078.004 plus T1059.009, and the agent hunt says so in its own header rather than inventing an identifier that would look tidier on a slide. That refusal is one of the thirteen counted further up this section.

Where this comes from Every identifier is quoted from the header comment of the file it sits in, and from the table headed MITRE ATT&CK in the two directory READMEs, which together carry 13 identifier rows and 12 distinct techniques. The defect behind each hunt is written in the file, above the query. These three files carry 0 tests, and the repository says so rather than letting a green suite badge imply otherwise. They are Kusto: running one needs a live workspace, which would break both the no dependency rule and the no network rule. No detection rate, no true positive rate and no tuning figure is claimed for any of them, because none has been measured against a real tenant that could be published.

Figures on this page were last re-derived from a run on 2026-09-20: the whole suite, the per module counts through tests/check_claims.py (opens the file on GitHub in a new tab), and all 137 mutations through tests/mutation_harness.py (opens the file on GitHub in a new tab). Both run in CI on every push. If any figure here stops reconciling with its parts, this page fails to build rather than rendering the stale number.

Frameworks and standards I work within
NIST CSF 2.0NIST AI RMF 1.0ISO/IEC 42001EU AI ActHIPAASOC 2FISMAMITRE ATT&CKMITRE ATLASOWASP LLM Top 10Zero Trust
What each one means, and how I use it

Plain-language, no jargon. The frameworks above are the standards that keep AI safe, legal, and accountable. Here is what the core ones are and how they show up in my work.

NIST AI RMF 1.0

The US government playbook for managing AI risk. It splits the job into four plain steps: set the rules (Govern), understand each system (Map), test it (Measure), and fix what you find (Manage).

How I use it: I use it as the backbone of the program, so every AI system gets an owner, a risk rating, and a review before it goes live.
In practice: Built an AI system inventory, risk-rated each tool, and put a documented approval gate in front of any new deployment.

ISO/IEC 42001

The first international, certifiable standard for running AI responsibly. Think of it as the AI version of a quality stamp: proof the process is repeatable and auditable, not improvised.

How I use it: I structure the governance documents and controls to its clauses, so the program could stand up to a formal audit.
In practice: Authored the AI acceptable-use policy and mapped each control to a clause, so the evidence is ready for an auditor.

EU AI Act

Europe AI law. It sorts AI into risk tiers (banned, high risk, limited, minimal) and puts the strictest rules on high-risk uses like healthcare diagnosis.

How I use it: I map each AI use case to its tier, so we know which legal obligations apply before deployment, not after.
In practice: Sorted each use case into a risk tier and routed the high-risk, clinical ones through human-in-the-loop approval.

HHS AI strategy

The US health department direction for using AI safely in healthcare, the sector where this work actually lives.

How I use it: I keep adoption guidance aligned to it so the program fits a regulated, healthcare environment from day one.
In practice: Routed regulated data to on-device, air-gapped inference so protected health information never leaves the boundary.

OWASP LLM Top 10

The ten most common security risks specific to AI chat and agent apps, such as prompt injection and sensitive-data leakage.

How I use it: I design guardrails against each one: input filtering, least-privilege tools, and checks on what the model sends back.
In practice: Set up content filters and prompt-injection guardrails in Azure AI Foundry, and wrote Microsoft Purview DLP policies so sensitive data cannot leave in a model response.

MITRE ATLAS

A catalog of real attacks against AI systems, the AI counterpart to the well-known ATT&CK matrix that security teams already trust.

How I use it: I threat-model agents against it to find how an attacker would actually try to break them, then close those paths.
In practice: Threat-modeled agents against known AI attack techniques, then added tool permissioning and egress secret masking to shut the gaps.

Credentials

Certifications

10 current certifications spanning security operations, architecture, AI, identity, and cloud. Listed below alongside 2 expired and open to renewal and 3 taken or booked but not yet conferred, each labelled with its status, plus a focused 2026 roadmap into offensive-aware architecture and advanced AI. Domains covered:

Security OperationsCybersecurity ArchitectureAI EngineeringIdentity and AccessCloud ArchitectureAI Governance
Certifications and status
Security+
CompTIA Security+
ExpiredCompTIA
Expired, open to renewing on request
SC-900
Security, Compliance, and Identity Fundamentals
CurrentMicrosoft
Credential ID: 56E64857DE863DAA
SC-200
Security Operations Analyst
ExpiredMicrosoft
Expired, renewing soon
SC-100
Microsoft Cybersecurity Architect
CurrentMicrosoft
Credential ID: 27F55F626C5DF2CB
Renewed August 2026 to August 2027
AZ-104
Azure Administrator Associate
CurrentMicrosoft
Credential ID: 3EA3312EE7C78D28
August 2026 to August 2027
AI-900
Azure AI Fundamentals
CurrentMicrosoft
Credential ID: EE8EEAC6A87AFA77
AI-102
Azure AI Engineer Associate
CurrentMicrosoft
Credential ID: 3C42BF552D8C4AD7
AI-103
Azure AI Apps and Agents Developer Associate
CurrentMicrosoft
Credential ID: 9DF77BD803E6940
AB-100
Agentic AI Business Solutions Architect
CurrentMicrosoft
Credential ID: 4DF3367C1813B11D
AB-741
AI Transformation Leader
CurrentMicrosoft
Credential ID: CE213C3CD43CCE3E
AZ-305
Azure Solutions Architect Expert
CurrentMicrosoft
Credential ID: 3C7479B53C192AF0
June 2026 to June 2027
SC-500
Cloud and AI Security Engineer Associate
CurrentMicrosoft
Credential ID: 694C88BB5814B051
AI-500
Multi-Agent AI Solutions Expert
Not yet conferredMicrosoft
Beta exam taken, awaiting scoring
AI-200
Azure AI Cloud Developer Associate
Not yet conferredMicrosoft
Booked, October 2026
CISSP
Information Systems Security Professional
Not yet conferredISC2
Booked, September 2026
2026 roadmap
AI-300Early October 2026
Microsoft AI
PenTest+October 2026
CompTIA PenTest+
OSCPNovember 2026
OffSec Certified Professional

Every credential below carries its issuing body and its credential identifier, so any of them can be checked directly with the issuer. Status is shown separately on each card.

By the numbers

The toolkit, connected

Explore the technology inventory, compare domains, and trace the certification timeline. Counts come from the listed entries, with credential status shown separately from technology coverage. Each chart includes a readable data table and its counting method.

56
distinct technologies named

grouped into 7 domains

15
certifications on the list

10 current, 2 expired, 3 not yet conferred

71
documented items in total

technologies plus certifications

Interactive field guide

Explore the toolkit

Select a domain, search a technology, and see how the toolkit fits together.

56
technologies
7
domains

Choose a domain to inspect its technologies.

Selected domain

AI and agentic tooling

Model providers, agent frameworks, and the editors and runtimes that host them, cloud and local.

13 listed technologies01 / Inspect
  • Claude Code
  • OpenAI
  • Azure AI Foundry
  • Cursor
  • Antigravity
  • Gemini
  • Ollama
  • LM Studio
  • Microsoft Copilot Studio
  • Microsoft Copilot
  • Hermes
  • Kimi
  • Agentic processes

Counts come from the listed technology inventory. Bar lengths use one shared scale. Orbit positions are decorative, not proficiency scores or relationships.

Documented coverage by domain

How much material the resume actually puts behind each domain, counted in items. The dashed inner shape is technologies. The hatched band out to the solid line is the certifications that evidence the same domain.

Explore the evidence
  • Technologies listed, dashed edge
  • Certifications on top, hatched, solid edge

Documented coverage by domain, counted in items. AI and agentic tooling: 13 technologies plus 7 certifications, 20 items. Security operations and exposure: 9 technologies plus 5 certifications, 14 items. Automation, data, and engineering: 8 technologies plus 0 certifications, 8 items. Endpoint, network, and infrastructure: 8 technologies plus 0 certifications, 8 items. Cloud and workplace platforms: 7 technologies plus 2 certifications, 9 items. Service management and operations: 6 technologies plus 0 certifications, 6 items. Identity and access: 5 technologies plus 1 certification, 6 items.

How this is counted Coverage for a domain is a plain sum: technologies listed in that domain plus certifications that evidence it. Both are counts of real resume entries, so the seven axes sum to 71, which is the 56 distinct technologies plus the 15 certifications. Each certification is assigned to exactly one domain. The scale is in items, not percent. This is not a proficiency score, a rating, or a self assessment. No such number exists in the resume, so none is plotted here. A short axis means the resume names fewer items under that heading, and nothing more than that.

Show the coverage counts
Coverage is technologies plus certifications, per domain. Every figure is a count.
DomainTechnologiesCertificationsCoverage (items)
AI and agentic tooling13720
Security operations and exposure9514
Cloud and workplace platforms729
Automation, data, and engineering808
Endpoint, network, and infrastructure808
Service management and operations606
Identity and access516
Total561571

Where the credentials come from, what they cover, where they stand

All 15 credentials, traced through three stages. Each one enters at one issuer, covers one domain, and ends at one status, so every column adds to 15 and every ribbon is a whole number of credentials.

Issuer
  • Microsoft13 of 15
  • CompTIA1 of 15
  • ISC21 of 15
Domain it evidences
  • AI and agents7 of 15
  • Security ops5 of 15
  • Cloud2 of 15
  • Identity1 of 15
Status today
  • Current10 of 15
  • Expired2 of 15
  • Not yet conferred3 of 15

13 of 15 come from one issuer, Microsoft. The stack is deep in one vendor ecosystem rather than spread thin across many.

7 of 15 sit on one domain, AI and agents, which is also the largest domain in the technology inventory above.

10 of 15 are current today. The rest are the two expired and the three taken or booked and awaiting conferral.

Worth saying out loud, because the diagram shows it as an absence: 3 of the 7 domains in the inventory carry no certification at all (Automation, Infrastructure, Service mgmt). Those are tooling and practice rather than paper, and the chart is not hiding that.

A flow diagram of 15 credentials in three stages. By issuer: Microsoft 13, CompTIA 1, ISC2 1. By domain: AI and agents 7, Security ops 5, Cloud 2, Identity 1. By status: Current 10, Expired 2, Not yet conferred 3.

How this is counted One credential is one unit of ribbon width. Every credential on the list is routed to exactly one issuer, one domain, and one status, so the three columns each total 15 and nothing can be counted twice or lost. Issuer and status are read from the site’s corrected credential list. The domain a credential evidences is the editorial part, using the same seven domains as the inventory above; the routing for every credential is printed in the table, so any assignment can be checked or disagreed with. No credential is weighted, scored, or ranked: a ribbon is thick because it carries more credentials, not because anything was judged more important.

Show the routing for every credential
Every credential, and the issuer, domain, and status it is routed through.
CredentialIssuerDomain it evidencesStatus
Security+, CompTIA Security+CompTIASecurity opsExpired
SC-900, Security, Compliance, and Identity FundamentalsMicrosoftIdentityCurrent
SC-200, Security Operations AnalystMicrosoftSecurity opsExpired
SC-100, Microsoft Cybersecurity ArchitectMicrosoftSecurity opsCurrent
AZ-104, Azure Administrator AssociateMicrosoftCloudCurrent
AI-900, Azure AI FundamentalsMicrosoftAI and agentsCurrent
AI-102, Azure AI Engineer AssociateMicrosoftAI and agentsCurrent
AI-103, Azure AI Apps and Agents Developer AssociateMicrosoftAI and agentsCurrent
AB-100, Agentic AI Business Solutions ArchitectMicrosoftAI and agentsCurrent
AB-741, AI Transformation LeaderMicrosoftAI and agentsCurrent
AZ-305, Azure Solutions Architect ExpertMicrosoftCloudCurrent
SC-500, Cloud and AI Security Engineer AssociateMicrosoftSecurity opsCurrent
AI-500, Multi-Agent AI Solutions ExpertMicrosoftAI and agentsNot yet conferred
AI-200, Azure AI Cloud Developer AssociateMicrosoftAI and agentsNot yet conferred
CISSP, Information Systems Security ProfessionalISC2Security opsNot yet conferred

Contact

Let us talk

I am open to conversations about agentic AI security, cloud security architecture, and roles where the two meet. If you are fully interested, please email me at ulisesghurtado@gmail.com.