Skip to content

Ingest Pipelines

Raw log messages are unstructured text — searching them is like reading a book without an index. Ingest pipelines automatically extract structured fields (IP addresses, status codes, request IDs, latencies, error levels) at index time, turning every log line into a queryable document. Filter by HTTP 500s, sort by response time, or aggregate errors by service without writing a single parser.

Ingest pipeline

How they work

Each subscription stream can specify a pipeline name (for example lambda, nginx, or json). Parsing runs server-side on the OpenSearch cluster during indexing — no Lambda processing overhead, no additional cost, and no code to maintain.

The raw message is always preserved alongside the extracted fields, so you never lose data. If parsing fails (malformed input or unexpected format), the event is indexed unchanged — there is no data loss.

Pipelines only affect OpenSearch; datalake writes remain raw for Athena query-time parsing.

14 built-in pipelines

The built-in pipelines cover the most common AWS and application log formats:

Pipeline Pipeline
Lambda VPC Flow Logs
JSON ALB
Nginx API Gateway
Apache RDS slow queries
Syslog EKS / Kubernetes
Tomcat CloudFront
Spring Boot RDS PostgreSQL

Custom pipelines

Create custom pipelines for proprietary formats via the OpenSearch Dev Tools console (the OpenSearch API). They become immediately available to any subscription stream.