
Build an API spec out of the traffic you already have.
openapi-enrich enriches an OpenAPI 3.1
document from observed HTTP traffic. Feed it a document and a set of recorded
request/response pairs, and it adds the paths, operations, parameters, request
bodies, and response schemas it can infer from them.
Introduction
Plenty of APIs have no specification, or one that stopped matching reality some time
ago. What they do have is traffic. This module treats that traffic as the source of
truth: record some real calls, and get a document describing what the API actually
does.
That document is also the first step towards a client library. It feeds
openapi-flatten,
openapi-compress and
openapi-codegen, so an API
that was never properly documented gets both a specification and a Go library to
call it with. Where the library fails to decode a response, recording that call
and enriching again closes the gap.
Enrichment is incremental by design. Every additional interaction refines the
result rather than replacing it — a second observation of the same endpoint
contributes any fields the first one didn't include, and widens a type where the two
disagree. That merging is delegated to
openapi-merge, which exists
precisely for the problem of reconciling schemas inferred from independent samples.
Features
What it infers:
- Paths — detected from request URLs, with ID-like segments replaced by
{param} path parameters.
- Operations — one per unique method + path, with an inferred
operationId
(e.g. GET /users → ListUsers, GET /users/{id} → GetUserByID).
- Query parameters — schema inferred from values; comma-separated values
become non-exploded arrays.
- Request headers —
Authorization creates an HTTP security scheme;
x-* and other custom headers become header parameters.
- Security — an operation only called without
Authorization gets
security: [], and one called both with and without it gets the credential
as optional ({} beside the scheme). A security list every operation shares
is stated once at the document level.
- Request bodies — JSON bodies produce inline object schemas.
- Responses — JSON, text/plain, and text/html responses are modeled;
repeated observations are merged.
- Arrays of objects — the elements of a recorded array of objects meet the
specification one by one, so in a list of mixed variants, such as Notion's
blocks, each reaches the variant of a union it matches rather than all of them
one. With no union there, they merge into one item as before.
- Binary bodies — a request or response body that is not text, such as a
zip, a PDF, an image or a video, is documented by its media type as a string
of bytes (
{"type": "string", "format": "binary"}), from its headers alone.
- Schema formats — UUID, URI, email, date-time, date, IPv4, IPv6 are detected
automatically from string values.
- Nulls and empty arrays — a value only ever seen as
null has the type
null, and becomes nullable once it is seen with a real type
(["string", "null"]). An array only ever seen empty is
{"type": "array", "maxItems": 0}, until a non-empty one shows its items.
- Schema types — a schema in the given document that has no
type gets
the one its enum or const values share, e.g. {"const": 401} becomes
an integer; a null among them makes it nullable.
- Enums — an
enum already declared in the given document grows with
every value observed for it. An object's keys count too, when its
propertyNames declares an enum. A recording never starts an enum of its own.
- Shared components — a schema several operations refer to is documented
from whichever of them was recorded, since sharing says they have the same
shape. Recording the others widens it to fit all of them. Where the sharing
itself is wrong, give each operation its own schema in the input.
The module also ships the pieces needed to obtain that traffic:
cassette — self-contained HTTP interaction types, with JSON persistence,
bearer-token masking, and header trimming before anything is written to disk.
A body that is not text is never read or written, only marked bodyOmitted,
so a large download streams to its caller as it is. A JSON string longer than
2,048 bytes, such as an image in base64, is cut to that and ends in ….
recorder — an http.RoundTripper that records live traffic into a cassette,
so you can capture interactions by pointing an existing client at it.
Usage
go get -tool github.com/MarkRosemaker/openapi-enrich/cmd/openapi-enrich
or
go get github.com/MarkRosemaker/openapi-enrich
import (
enrich "github.com/MarkRosemaker/openapi-enrich"
"github.com/MarkRosemaker/openapi-enrich/cassette"
)
// Start from a minimal document or load an existing spec.
doc := enrich.NewDocument()
interactions := []cassette.Interaction{
{
Request: cassette.Request{
Method: "GET",
URL: "https://api.example.com/users",
Headers: http.Header{},
},
Response: cassette.Response{
StatusCode: http.StatusOK,
Headers: http.Header{"Content-Type": {"application/json"}},
Body: []byte(`[{"id":1,"name":"Alice"}]`),
},
},
}
if err := enrich.Enrich(doc, interactions); err != nil {
log.Fatal(err)
}
The main public function is:
func Enrich(doc *openapi.Document, interactions cassette.Interactions) error
Schemas are left inline — the caller composes any post-processing as needed.
Design
- No I/O — the caller loads and saves the spec.
- No flatten/tidy/sort — use separate libraries for those.
- Own interaction types — no dependency on a specific HTTP recording format.
The result of enrichment is deliberately raw: inline schemas, unsorted, unpolished.
Turning that into something pleasant to read or generate from is the job of the
modules below, applied in whatever order suits you.
The openapi family
| Module |
Purpose |
| openapi |
Parse, validate, and write OpenAPI 3.x specifications |
| openapi-compare |
Compare specification objects — exact equality and shape equivalence |
| openapi-edit |
Safe structural edits, such as renaming a schema and rewriting every $ref to it |
| openapi-flatten |
Promote inline definitions into named components entries |
| openapi-compress |
Deduplicate and merge equivalent component schemas |
| openapi-merge |
Merge schemas that were inferred independently from different samples |
| openapi-enrich (this module) |
Infer specification content from observed HTTP traffic |
| openapi-codegen |
Generate Go types, clients, and servers from a specification |
A common sequence is to enrich from traffic, flatten the inline schemas into named
components, compress the duplicates that flattening produces, and then generate a
client.
Contributing
Contributions are welcome — please open an issue or a pull request on GitHub.
License
This project is licensed under the Apache 2.0 License.