Get started

Security graph schema

Schema v0 is a closed vocabulary. Collectors, the plugin SDK, and the API all use the same type strings. Unknown types are rejected when a plugin batch is validated. New attributes go in properties so existing rows keep working.

Tables

om migrate creates nodes and edges. Both rows carry properties JSONB. Edges reference nodes with ON DELETE CASCADE. The unique key on an edge is (source_id, target_id, type).

Node columns: id, type, name, provider, region, account_id, properties, timestamps.

Node types

Type Role
Internet Synthetic entry point. Id is always internet:global.
Network Security group, Kubernetes namespace, or Service.
Workload Compute: instance, VM, or pod.
Identity Principal that can be assumed or granted access.
Datastore Object or account storage.
Finding CSPM, CVE, exposure, or attack-path result. finding_type is cspm, cve, exposure, or attack_path.
Control Reserved guardrail node. The Kubernetes collector writes the cluster as a control.

Edge types

Type Typical direction
REACHABLE internet:global → workload
ASSUMES Workload → identity on AWS. Service account → pod on Kubernetes.
CAN_ACCESS Identity or workload → datastore
AFFECTS Workload → network control
VIOLATES Finding → affected resource

IDs

Built-in collectors use these patterns:

Provider Format
AWS aws:{account}:{region|global}:{kind}:{resource}
Azure azure:{subscription}:{location}:{kind}:{resource}
GCP gcp:{project}:{region-or-zone}:{kind}:{resource}
Kubernetes k8s:{cluster}:{kind}:{resource}
Finding finding:{rule-id}:{resource-node-id}. An attack-path finding is finding:attack-path:{source-finding-id}:{datastore-id}.
Edge {source}|{target}|{type}

kind is a short word such as workload, identity, or datastore, not the node type string. AWS uses the scan region for EC2 and security groups, and global for IAM and S3. Azure’s location segment is AZURE_LOCATION, not the resource region stored on the node. GCP instances use the zone. GCS buckets and service accounts use GCP_REGION. Plugins may use any unique id. Every edge endpoint in a plugin batch must be a node in that same batch.

Properties rules and queries read

Anything else may be stored. These keys change rule matches and named queries:

  • public_access (bool) — datastore is treated as public
  • sensitivity (string) — crown-jewel mark on a datastore. Any non-empty value is enough. internet-to-sensitive-datastore keeps those paths, and internet-to-datastore does not filter on it. Collectors copy a tag or label named sensitivity or data-class (sensitivity wins). A blank value is omitted. A plugin may set the property directly. Object contents are not read.
  • public_access_block — disabled matches the public-bucket query and a CIS-inspired rule
  • encryption, versioning, service — S3 pack rules
  • open_ingress — security group allows 0.0.0.0/0 or ::/0
  • admin_access (bool) — identity is treated as administrative
  • mfa, unused_access_keys — IAM user pack rules
  • public_ip, imdsv2 — workload exposure and metadata rules
  • packages, image, images — workload inventory that om enrich cve matches. Rules do not read these keys. See CVE enrichment.
  • audit_events — recent CloudTrail, Activity Log, or Admin Activity events on an identity or resource. Each item has id, name, time, principal, resource, principal_node_id, resource_node_id, and optional source_ip and read_only. om scan aws, om scan azure, and om scan gcp write at most five, newest first, and only when the resource is on an exposed path. A plugin may set the same list. Named path queries copy an event into audits when its resource node is on the path. Rules do not match this key.
  • internet_reachable, path_to_datastore, admin_can_access — graph match keys on YAML rules, computed at run time, not stored by collectors

The design record is ADR 001. Plugin authors should use the constants in sdk/collector.