UUIDs and random IDs: v4 vs v7, collision math, and why Math.random is wrong

UUID versions explained, the real collision odds for v4, why Math.random is unsafe for IDs, and how ULID and nanoid compare.

Published 2026-09-25

What a UUID actually is

A UUID (Universally Unique Identifier) is a 128-bit value, almost always written as 32 hex digits split by hyphens into five groups: xxxxxxxx-xxxx-Mxxx-Nxxx-xxxxxxxxxxxx. The M position encodes the version (which algorithm generated it) and the N position's leading bits encode the variant (which layout rules it follows). The current specification is RFC 9562 (May 2024), which formally obsoletes the older RFC 4122 and adds several new versions — most notably v6 and v7 — on top of the original layout.

UUIDv4:  f47ac10b-58cc-4372-a567-0e02b2c3d479
                       ^--- version 4 here

v4 (random) vs v1/v7 (time-based)

The versions differ in what goes into those 128 bits:

  • v1 combines a 60-bit timestamp with the generating machine's MAC address. It's effectively unique by construction, but embedding a MAC address leaks information about the machine that created it — a real privacy concern that pushed later versions away from this design.
  • v4 fills the space with randomness: per RFC 9562, a v4 UUID sets the fixed version/variant bits and fills the remaining 122 bits with random data, generated from a cryptographically secure pseudorandom number generator (CSPRNG). This is the version most tools default to, precisely because it needs no coordination and leaks no identifying information.
  • v7 puts a 48-bit Unix millisecond timestamp in the most significant bits, then fills the rest with random data. Because it starts with a timestamp, v7 UUIDs generated later always sort after ones generated earlier by plain byte comparison — v4 UUIDs, being fully random, sort in no meaningful order at all.

The real collision math for v4

"UUIDs never collide" is a rough approximation, not a guarantee — but the actual odds are worth knowing rather than assuming. With 122 random bits, the birthday-bound math (the same math behind hash-collision odds) says you'd need to generate roughly 2.71 quintillion random v4 UUIDs before the probability of a single collision reaches 50%, per the calculation in the collision-probability discussion on Wikipedia's UUID article — equivalent to generating a billion UUIDs per second, continuously, for about 86 years. For essentially any application, that risk is negligible compared to other failure modes in the system. The one place it stops being negligible is if the source of randomness itself is weak — the math assumes genuinely random, independent 122-bit values, which is exactly where Math.random() breaks the assumption.

Why Math.random() is the wrong tool for IDs

Math.random() is a fast, general-purpose pseudorandom number generator, not a cryptographically secure one — its seeding and internal state are not designed to resist prediction, and in some JavaScript engines its output has been shown to be predictable from a handful of observed values. Using it to build a UUID-shaped string undermines the entire basis for the collision math above: the 2.71-quintillion figure assumes uniform, unpredictable 122-bit randomness, and a weak PRNG doesn't deliver that, whether or not the resulting string still looks like a UUID.

The browser and Node.js both ship the correct primitive: crypto.getRandomValues(), seeded from OS-level entropy sources rather than a simple internal PRNG state, which is what makes its output resistant to prediction. Most runtimes also now expose crypto.randomUUID() directly, which handles the version/variant bits correctly for you (in browsers it is only available in secure contexts, i.e. HTTPS or localhost):

// wrong: predictable, not suitable for IDs used as security tokens
const fakeId = Math.random().toString(36).slice(2);

// right
const id = crypto.randomUUID(); // e.g. "f47ac10b-58cc-4372-a567-0e02b2c3d479"

This hub's UUID generator and password generator both draw from a CSPRNG for exactly this reason — an ID or a secret that needs to be unguessable is only as strong as the randomness underneath it, and Math.random() was never built to provide that.

Why v7 matters for database indexes

This is the practical reason v7 exists, beyond sortability being a nice-to-have. A database index (typically a B-tree) stores rows in key order. Inserting fully random v4 UUIDs as primary keys means each new row lands at a random position in the index rather than at the end — which causes frequent page splits, worse cache locality, and index fragmentation under heavy insert load, compared to a monotonically increasing key. RFC 9562 says this directly: UUIDs that are not time ordered, such as v4, "have poor database-index locality" because new values are not close to each other in the index, and time-ordered v7 values insert near the end of the index, next to recently inserted ones, the same access pattern that made auto-incrementing integer keys efficient — while still giving you a 128-bit, effectively collision-resistant identifier with no central coordination, which auto-increment integers can't provide across distributed systems.

ULID and nanoid: the non-UUID alternatives

Two ID formats outside the UUID spec solve overlapping problems differently:

ULID is 128 bits like a UUID — a 48-bit millisecond timestamp plus 80 bits of randomness — but encodes as 26 characters of Crockford's Base32 instead of 36 characters of hyphenated hex, and per the ULID spec is lexicographically sortable by construction, the same underlying idea later standardized as UUIDv7. It's a reasonable choice if you want time-ordering and a shorter, URL-friendly string, and don't need strict UUID-format compatibility with existing tooling.

nanoid takes a different approach: it doesn't encode a timestamp at all, just a configurable number of random characters from a configurable alphabet (21 characters by default, using a URL-safe alphabet), aimed at being shorter than a UUID while keeping comparable collision resistance for that length. It's a good fit when you want the smallest reasonable random ID and don't need — or specifically don't want — the created-at information a timestamp-based ID leaks.

Which to use is mostly about what property you need: UUIDv4 for a no-coordination unique identifier with universal tooling support; UUIDv7 or ULID when insert-heavy database performance and natural time-ordering matter; nanoid when string length is the binding constraint and you don't want a timestamp embedded in the ID.

A practical checklist

  1. Default to v4 (or v7 if the ID becomes a database primary key under heavy write load) — this hub's UUID generator produces v4; for v7 use your language's UUID library.
  2. Never generate an ID or token with Math.random() if it needs to be unguessable — use crypto.randomUUID() or crypto.getRandomValues().
  3. If insert performance on an indexed primary key matters, prefer v7 or ULID over v4 — random keys fragment B-tree indexes under load.
  4. Don't use v1 for anything public-facing — it embeds the generating machine's MAC address in the ID.

Tools mentioned in this guide