Ingest Pipelines¶
Raw log messages are unstructured text — searching them is like reading a book without an index. Ingest pipelines automatically extract structured fields (IP addresses, status codes, request IDs, latencies, error levels) at index time, turning every log line into a queryable document. Filter by HTTP 500s, sort by response time, or aggregate errors by service without writing a single parser.

How they work¶
Each subscription stream can specify a pipeline name (for example lambda, nginx, or json). Parsing runs server-side on the OpenSearch cluster during indexing — no Lambda processing overhead, no additional cost, and no code to maintain.
The raw message is always preserved alongside the extracted fields, so you never lose data. If parsing fails (malformed input or unexpected format), the event is indexed unchanged — there is no data loss.
Pipelines only affect OpenSearch; datalake writes remain raw for Athena query-time parsing.
14 built-in pipelines¶
The built-in pipelines cover the most common AWS and application log formats:
| Pipeline | Pipeline |
|---|---|
| Lambda | VPC Flow Logs |
| JSON | ALB |
| Nginx | API Gateway |
| Apache | RDS slow queries |
| Syslog | EKS / Kubernetes |
| Tomcat | CloudFront |
| Spring Boot | RDS PostgreSQL |
Custom pipelines¶
Create custom pipelines for proprietary formats via the OpenSearch Dev Tools console (the OpenSearch API). They become immediately available to any subscription stream.