Skip to content

OSV Schema

The core type models the OSV Schema (currently 1.4.0).

Top-level structure

Required vs optional

FieldRequiredNotes
schema_versionCurrently 1.4.0
idUnique record identifier
modifiedLast modification time
publishedFirst publication time
withdrawnString, not time.Time
aliasese.g. CVE-2024-XXXX
affectedBut usually present
severityCVSS v2 / v3 / v4

osv validate enforces id and schema_version.

The check is intentionally shallow — it confirms the record is parseable and carries the two identity fields, not that every optional field is well-formed. There are three independent failure layers, checked in order: file readability, raw JSON syntax, then OSV struct decode. json.Valid passing does not guarantee UnmarshalFromJson succeeds — a record where, say, affected is a string instead of an array is syntactically valid JSON but fails OSV decoding and is reported as a parse error. Once decoded, id and schema_version are checked as two independent ifs, not a short-circuit: both errors are collected when both are empty (a record missing only id still gets checked for schema_version). affected, severity, and references are not checked; a record with no affected entries still validates.

Full type relationship

Affected → package → ranges → events

The package object carries three fields: ecosystem (one of the typed constants), name (the package name — for Maven this is groupId:artifactId), and purl (an optional Package URL string). purl is informational; the SDK doesn't parse it, so for ecosystem-specific decomposition (like Maven GAV) use name via GetGroupID / GetArtifactID, not purl. An affected entry may also carry its own severity slice, scoped to that affected range.

Lifecycle of a record

Field quick-lookup by intent

Is a version affected? — event-timeline resolution

The single most important algorithm when consuming OSV data is: given a concrete version, is it vulnerable? OSV answers this not with prose but with the ordered events inside each range. You walk the timeline left to right, toggling an "affected" flag.

The special value introduced: "0" means "from the very first version". The flow above covers the three common event kinds (introduced / fixed / last_affected); the fourth, limit, marks a range's upper bound and is not the same as last_affectedlimit is exclusive, so reaching V >= limit clears the flag (the limit version itself is not affected), whereas last_affected is inclusive (the last_affected version is affected, only V > last_affected clears). limit is rare outside GIT ranges. The SDK gives you the per-event predicates to implement this yourself:

Why the golden rule matters here

Because each event carries exactly one non-empty key, the walk above can switch on "which predicate is true" without ambiguity. That is also why osv query --events emits omitempty JSON — a stray "fixed": "" would make two predicates look true.

RangeType — how versions are compared

The < / >= comparisons in the algorithm above are not universal string comparisons. The range's type decides the ordering rules.

RangeTypeConstantVersion tokens are…
SEMVERRangeTypeSemverSemVer 2.0.0 strings, compared by precedence
ECOSYSTEMRangeTypeEcosystemOpaque strings ordered by the ecosystem (PyPI→PEP 440, etc.)
GITRangeTypeGitGit commit hashes, resolved via the commit graph

GIT ranges are not string-sortable

For GIT ranges you cannot decide affectedness by comparing hash strings — you need the repository's commit ancestry. Treat GIT ranges as "requires graph resolution", not "compare like SEMVER". The Range.Repo field (the repo URL) is what anchors a GIT range to that ancestry; for SEMVER / ECOSYSTEM ranges it is usually empty.

Severity scoring internals

severity[].score holds a CVSS vector string, not a number. The SDK exposes three getters that share one lazily-parsed, memoized backing value.

GetterOn a vector stringUse when
GetScore()0.0You just want a float and treat 0 as "n/a"
GetScoreAsFloat()(0, error)You must distinguish a real 0 from a parse failure
GetScoreAsPointer()nilYou want nil to mean "no numeric score"

To rank severity when the score is a vector, read SeveritySlice.GetCVSS3() / GetCVSS2() and interpret the vector — see Skills → severity.

Serialization: one struct, six tag namespaces

Every core field is tagged for six ecosystems at once, so the same struct round-trips through JSON, YAML, config decoding, raw SQL, MongoDB, and GORM without adapters.

Database strategy: columns vs JSON blobs

Simple scalar fields become plain columns. Complex nested slices (AffectedSlice, SeveritySlice, Range, …) implement sql.Scanner + driver.Valuer, so GORM stores them as a single JSON string and rehydrates them on read.

Generic type parameters

OsvSchema[EcosystemSpecific, DatabaseSpecific] carries two type parameters that flow down into Affected and Range, so vendor-specific blobs stay typed instead of collapsing to map[string]any.

For everyday parsing use [any, any] (as every CLI command does). Supply concrete structs only when you need typed access to ecosystem_specific / database_specific.

Source files

All types live in the root package osv_schema:

FileContents
osv_schema.goOsvSchema top-level type
package.goPackage, Ecosystem constants
affected.goAffected, AffectedSlice
severity.goSeverity, SeveritySlice
range.goRange
event.goEvent
references.goReferences
aliases.goAliases
related.goRelated
credits.goCredits
unmarshal.goUnmarshalFromJson / UnmarshalFromJsonFile

Last updated:

Released under the MIT License.