Get the Zoom link
Cloud Security Office Hours Banner

Cloud security how-to guides

The small languages almost everything practical runs on: policy, query, pattern, and detection. Each guide starts with what the thing actually is, walks through it hands-on, and finishes with exercises and worked answers.

Browse the guides Which one first?

· · Vendor-neutral

Cloud security has a set of small languages sitting underneath almost everything practical. A policy that blocks a deploy is Rego or CEL. A detection that fires in your SIEM is Sigma compiled into something else. A vulnerability feed is OVAL. The command that answers "which of our roles can do this" is jq or JMESPath. The document that decides whether a request succeeds is IAM JSON or Cedar.

None of them is large. Each can be learned properly in an afternoon, and each is the kind of thing people instead learn by copying a snippet that works, changing it until it stops working, and never quite finding out why. That gap is where the expensive mistakes live: a detection that silently never fires, an allowlist regex that a query string walks straight past, a policy whose typo turned it into decoration.

These guides all follow the same shape. An explanation of what the thing actually is and the one idea that makes it click, a walk-through you can copy and run, and hands-on exercises with worked answers. Every command and every rule on these pages was run before it was published, except on the Sumo Logic, Datadog and YARA-L guides, whose engines run only inside the vendors' services; those guides were checked against the vendors' documentation and say so at the top.

A theme runs through all of these, and it is worth naming up front. Nearly every failure mode in this collection is silent. A wrong Rego rule returns an empty deny set. A Sigma rule converted without a pipeline returns zero results. A JMESPath filter with double quotes returns an empty list. A mistyped IAM condition key produces a valid policy that never matches. None of these raise an error, and all of them look exactly like good news. The habit each guide keeps returning to is running a control: prove the instrument works on something you know it should catch, before believing what it tells you is absent.

On this page

  1. The guides
  2. Detection query languages
  3. Which one to learn first
  4. Choosing between the policy languages
  5. Where next

The guides

Regex for security work

Regex as a control rather than a search: anchors and allowlist bypasses, credential scanning and where it stops working, backtracking engines versus RE2, and catastrophic backtracking you can time for yourself.

BeginnerFoundationalLocal / $0~2h

jq and JMESPath

The two JSON query languages you will use daily, why they are not interchangeable, and the traps that return an empty result instead of an error. Includes a sample dataset, so no cloud account required.

BeginnerFoundationalLocal / $0~90m

IAM policy languages and Cedar

Why a policy that says Allow may grant nothing, the IAM elements that invert their apparent meaning, and Cedar, whose schema catches at write time the typo that IAM lets fail silently forever.

IntermediateAWSCedar~2h

OPA and Rego

The closest thing to a universal policy language. Undefined versus false, deny sets, gating a real Terraform plan with Conftest, and unit-testing the policy in both directions.

IntermediatePolicy as codeTerraform~90m

CEL policy expressions

The expression language embedded in the Kubernetes API server and Google Cloud IAM. Macros, the presence check that decides whether your policy fails open or closed, and a real admission policy rolled out safely.

IntermediateKubernetesGCP~90m

Sigma and YARA

Detection as code in two formats: Sigma for log events, compiled to your SIEM's dialect, and YARA for bytes in files and memory. Both tested against benign data as well as malicious.

IntermediateDetectionLocal / $0~2h

OVAL and SCAP

The machine-readable substrate under compliance scanning and vulnerability feeds. How to read a definition, and why backported patches make upstream version comparison wrong in the direction that wastes the most time.

IntermediateComplianceVuln management~2h

Detection query languages

These guides cover the languages detections are written in: the query language of each SIEM (security information and event management) platform or data store, and the rule formats around them. They are bigger than the languages above, and none of them assumes you already know SQL or any other query language. All of them use the same 24-event AWS CloudTrail lab, a leaked continuous-integration access key followed through the logs, and the same five detections, so once you have worked through one, the others read as translations and the differences between platforms stand out.

Most run on your own machine at no cost: DuckDB, a Splunk Enterprise trial, Microsoft's Kusto emulator, Elasticsearch, Panther's command-line tool and Falco. Sumo Logic, Datadog and Google Security Operations run only as hosted services, so those guides were checked against the vendors' documentation rather than run, and say so at the top.

SQL for security data lakes

Learn SQL from zero on 24 CloudTrail events in DuckDB, free on your own computer, including the NULL comparison that silently drops events. Then port the detections to Snowflake, BigQuery and Amazon Athena.

BeginnerDetectionLocal / $0~2h

Splunk SPL

Learn Splunk's search language from zero on 24 CloudTrail events in a local Splunk Enterprise: filter, count, bucket and correlate, then schedule an alert and test it both ways. Includes the searches that return nothing because they cannot work.

BeginnerDetectionLocal / $0~2h

Kusto Query Language (KQL)

Learn KQL from zero on a local Kusto engine and 24 CloudTrail events: filter, parse, count and join, then write a Sentinel analytics rule and test it both ways. Includes the rule that silently misses a burst.

BeginnerDetectionLocal / $0~2h

Elastic ES|QL and EQL

Query CloudTrail in Elasticsearch with ES|QL and EQL, from zero: counts, thresholds and sequences on a local lab, then a rule checked with Elastic's validator. Includes the silent failures that make rules return nothing.

IntermediateDetectionLocal / $0~2h

Sumo Logic search queries

Learn Sumo Logic's search language from zero on 24 CloudTrail events: parse, filter, count and correlate, then build a monitor and test it both ways. Includes the parse step that silently drops events.

BeginnerDetectionVendor tenant~2h

Datadog Cloud SIEM detection rules

Learn Datadog's log search and Cloud SIEM rule settings from zero on 24 CloudTrail events: filter, group, count, sequence and test both ways. Includes where Datadog puts CloudTrail's fields, and a threshold that silently misses EC2 refusals.

BeginnerDetectionVendor tenant~2h

YARA-L 2.0

Learn YARA-L 2.0 from zero on 24 CloudTrail events: UDM fields, placeholders, match windows and two-event joins, checked against Google's documentation. Includes the parser change that can silence a rule copied from Google's own library.

IntermediateDetectionVendor tenant~90m

Panther Python detections

Write Panther's Python rules from zero on 24 CloudTrail events and test them on your own computer with panther_analysis_tool. Includes the refusal count that silently comes up short, and which checks only a Panther deployment runs.

BeginnerDetectionLocal / $0~90m

Falco rules

Write Falco rules from zero: conditions, lists, macros, exceptions and overrides, run on a CloudTrail lab and then on a container's system calls. Includes why Falco cannot count, and where the counting goes instead.

IntermediateDetectionLocal / $0~2h

What changes from one to the next is mostly where counting and correlation live:

Guide Where the count lives Two events in order Practise it locally
SQLGROUP BY and HAVING in the queryA join of the table with itselfDuckDB
Splunk SPLstats in the searchtransactionSplunk Enterprise trial in Docker
KQLsummarize in the queryjoinKusto emulator in Docker
ES|QL and EQLSTATS in ES|QLEQL sequenceElasticsearch in Docker
Sumo Logiccount by in the querytransactionNo: hosted only
Datadog Cloud SIEMThe rule's settings, not the searchThe Sequence detection methodNo: hosted only
YARA-L 2.0A match window and a count in the conditionTwo event variables joined on a shared valueNo: needs a Google SecOps tenant
PantherThe rule's YAML file (Threshold)A correlation rule, evaluated only by a Panther deploymentPython rules with panther_analysis_tool
FalcoNot in Falco: wherever its alerts goNot in Falco: wherever its alerts goFalco in Docker

Which one to learn first

The cards above are in the order I would actually take them, which is not the order of how interesting they are. It is the order of how often each one is the thing standing between you and an answer.

  1. Regex, then jq and JMESPath. These two are the floor. Everything else on this page either contains them or assumes them: Sigma rules embed regex, YARA strings can be regex, CEL has a matches() function, and auditing IAM at scale is a jq problem. Time spent here pays back inside a week.
  2. IAM policy languages. If you work in cloud security and read one kind of document more than any other, it is this one. It is also the one where being approximately right is most dangerous, because approximately right produces a confident wrong answer about who can do what.
  3. Then a policy language, chosen by what you are enforcing. Rego if you are gating infrastructure code in CI. CEL if you are working in Kubernetes or Google Cloud. The section below is about making that choice deliberately.
  4. Sigma and YARA when detection is your job, and OVAL when compliance or vulnerability management is. Both are worth skimming even if they are not, so that you recognise what you are looking at when someone hands you a rule or a scan result.
  5. Then the detection language your stack runs. Each detection guide teaches its language from zero, and they share one lab and five detections, so the first one you take is the hard one. With no stack yet, start with SQL, which runs free on your own computer and carries over to data lakes and to Panther's scheduled searches.

If you are earlier in the journey than this, the learning path is a better starting point, and the home lab walk-throughs give you an environment to run any of this against.

Choosing between the policy languages

Rego, CEL, and Cedar overlap enough to be confusing and differ enough that picking the wrong one is expensive. The short version:

  Rego (OPA) CEL Cedar
Runs where A separate engine: sidecar, service, or CLI In-process, inside the host that is deciding A library, or Amazon Verified Permissions
Best at Arbitrary policy over arbitrary JSON Fast boolean checks on one object Who may do what to which resource
Can consult external data Yes, that is a core feature No, only what the host supplies Only the entity store you provide
Static validation opa check --strict, plus Regal Type-checked against the host's schema Full schema validation, including typos
Typical use Terraform and Kubernetes manifests in CI Kubernetes admission, GCP IAM conditions Application authorization
Main cost A component to deploy and keep available Cannot express anything needing a join Narrower scope by design

A rough decision rule that holds up in practice: if the question is "may this principal do this to this resource", reach for Cedar. If it is "is this configuration acceptable" and the answer depends only on the configuration itself, reach for CEL, especially where a host already embeds it. If the answer depends on anything outside the object being judged, that is Rego, and the deployment cost is the price of the capability.

Where next