# Skills & taxonomies

> The two classification vocabularies — how they're tokenised, how aliasing breaks round-tripping, and how to query them.

Source: https://hyperjobs.io/docs/guides/skills-and-taxonomies

Postings carry two independent classification vocabularies, and they answer
different questions:

- **`skills`** — specific, granular capabilities. `python`, `kubernetes`,
  `cpp`. About **84%** of postings have at least one (504,787 of 603,335).
- **`taxonomies`** — the job family. `Software`, `Healthcare`,
  `Finance & Accounting`. 28 fixed values, on about **64%** of postings
  (385,246).

They're orthogonal. A posting can be `Software` with no skills, carry
`kubernetes` with no taxonomy, or have both. Neither is guaranteed, which
matters more than it sounds: filtering on either one silently excludes the
postings that were never classified.

```bash
curl -G https://api.hyperjobs.io/v1/jobs \
  -H "Authorization: Bearer $HYPERJOBS_KEY" \
  --data-urlencode "skill=kubernetes" \
  --data-urlencode "taxonomies=Software"
```

## The token model

Skills are stored as **canonical lower-case tokens with no spaces and no
punctuation**. Not free text, not the words that appeared in the posting — a
normalised token drawn from a fixed vocabulary. `Machine Learning` in a job
description becomes the token `machinelearning`.

`skill=` is a token match, not a substring match. `skill=java` does not match
`javascript`, and `skill=py` matches nothing at all. Contrast with `title=` and
`description=`, which are Google-style searches — see
[filtering](/docs/guides/filtering).

## Two alias tables, pointing opposite ways

There are two separate translation layers, and conflating them is the source of
almost every "but I can see that skill in the response" bug.

<Fields>
  <ResponseField name="Query aliases (input)" type="7 tokens">
    Applied to the value you send in `?skill=`. Lets you write the punctuated
    form of a few well-known names and have it resolved to the stored token.
  </ResponseField>
  <ResponseField name="Display aliases (output)" type="one-way">
    Applied to the `skills` array on the way out, to render a token as something
    a human wants to read. **Has no reverse mapping.**
  </ResponseField>
</Fields>

The full query-alias table:

| You send | Resolves to | Rows |
| --- | --- | --- |
| `c++` | `cpp` | 9,586 |
| `c#` | `csharp` | 6,309 |
| `.net` | `dotnet` | 4,912 |
| `node.js`, `nodejs` | `node` | — |
| `react.js` | `react` | — |
| `vue.js` | `vue` | — |
| `golang` | `go` | — |

The punctuated forms are **not stored** — `c++` as a stored token is 0 rows, and
`c#` is 0. They work only because the input alias rewrites them before the query
runs. That's the whole table; anything not on it is sent through untranslated.

## The round-trip problem

<Warning title="A skill you can see in a response often won't work as a filter">
  Display aliases are one-way. There is no reverse mapping, so a value the API
  *showed* you is not necessarily a value the API will *accept*. Feed it back and
  you get zero rows — not an error, not a warning, just an empty set that looks
  exactly like "no jobs match".

  | You see | Actually stored as | `?skill=` with what you saw | `?skill=` with the token |
  | --- | --- | --- | --- |
  | `Machine Learning` | `machinelearning` | **0 rows** | 22,269 |
  | `CI/CD` | `cicd` | **0 rows** | 4,178 |
  | `Quality Assurance` | `qualityassurance` | **0 rows** | ✓ |
  | `C++` | `cpp` | ✓ — input-aliased | ✓ |

  Only the query-alias entries above survive the round trip. Everything else
  that gets display-formatted does not.

  **The rule:** if a displayed skill contains a space or a slash, strip them and
  lower-case it before filtering.

  ```js
  // Turns a displayed skill back into a queryable token.
  const toToken = (displayed) => displayed.replace(/[\s/]/g, "").toLowerCase();

  toToken("Machine Learning"); // → "machinelearning"
  toToken("CI/CD");            // → "cicd"
  ```

  This is a heuristic, not the inverse function — no such function exists. It
  covers the space-and-slash cases, which is the whole observed failure mode.
</Warning>

## Comma means AND — but only for `skill`

Comma means **OR** on every multi-value filter, with one exception:
`skill`, where **each value is another required condition**, not an
alternative.

```bash
# Tagged BOTH python AND aws. Not "either" — skill is the AND exception.
curl "…/v1/jobs?skill=python,aws"

# Either family — taxonomies is OR (overlap match).
curl "…/v1/jobs?taxonomies=Software,Healthcare"
```

`skill`'s AND is exactly right when you mean it: `skill=react,typescript,graphql`
finds roles that need all three, which is what a skill list usually means. Just
don't carry the habit elsewhere — every *other* comma list in the API reads as
OR now, and vice versa: a comma in `skill=` narrows where a comma anywhere else
widens.

### Getting OR across skills

`taxonomies` no longer needs this — `taxonomies=Software,Data & Analytics` is a
native union. `skill` still does: run one request per value and merge
client-side.

```js
const key = process.env.HYPERJOBS_KEY;

// The union of single-skill queries — the OR that `skill=` won't express.
async function anyOf(param, values) {
  const pages = await Promise.all(
    values.map((v) =>
      fetch(
        `https://api.hyperjobs.io/v1/jobs?${param}=${encodeURIComponent(v)}&limit=200`,
        { headers: { Authorization: `Bearer ${key}` } },
      ).then((r) => r.json()),
    ),
  );

  // Dedupe by id — a posting can legitimately carry several of the values.
  const byId = new Map(pages.flatMap((p) => p.data).map((j) => [j.id, j]));
  return [...byId.values()];
}

const rustOrGo = await anyOf("skill", ["rust", "go"]);
```

N values costs N requests against your [rate limit](/docs/rate-limits). To mix
AND with OR, put the AND terms in each request's comma list and vary the OR term
across requests — `skill=rust,kubernetes` and `skill=go,kubernetes` unions to
"(Rust or Go) and Kubernetes".

## The 28 taxonomies

Exact, Title-Case values. A casing miss or a value off this list is
a `400` naming the allowed values — never a silent zero-row response.

| | | |
| --- | --- | --- |
| `Software` | `Technology` | `Engineering` |
| `Data & Analytics` | `Healthcare` | `Finance & Accounting` |
| `Consulting` | `Legal` | `Human Resources` |
| `Marketing` | `Creative & Media` | `Sales` |
| `Art & Design` | `Education` | `Manufacturing` |
| `Energy` | `Construction` | `Science & Research` |
| `Customer Service & Support` | `Logistics` | `Administrative` |
| `Management & Leadership` | `Trades` | `Transportation` |
| `Hospitality` | `Security & Safety` | `Social Services` |
| `Agriculture` | | |

<Note>
  `Software` and `Technology` are separate families, as are `Software` and
  `Engineering`. If you're after technical roles broadly, ask for all three in
  one request — `?taxonomies=Software,Technology,Engineering` is a native OR
  (overlap match) — rather than assuming one subsumes the others.
  The full list also lives in the
  [enums reference](/docs/reference/enums#taxonomies).
</Note>

## Which one should you filter on?

| Use | When | Coverage | Cost |
| --- | --- | --- | --- |
| `taxonomies` | You want a whole job family — "all healthcare roles" | 64% | Coarse; misses the 36% unclassified |
| `skill` | You want a specific capability — "roles needing Rust" | 84% | Misses postings that never listed skills |
| `title` | You want a named role — "SRE", "Staff Engineer" | 100% | Google-style search; broad, still noisy |

The practical guidance:

- **Start with `skill`** for anything technology-shaped. It's the highest-signal
  filter and has the best coverage of the three classifiers.
- **Use `taxonomies` to scope, not to find.** It's good for narrowing a broad
  query to a sector and bad as the only filter, because a third of postings
  carry no taxonomy and are invisible to it.
- **Reach for `title` when the concept isn't a skill or a family.** "Staff" and
  "Principal" are seniority signals in the title that no vocabulary captures —
  though note `seniority` is its own exact-match filter.
- **Don't combine all three defensively.** Every classifier you add is another
  AND, and each one drops the postings it never classified. Three filters at 84%,
  64%, and 100% coverage intersect to a much smaller set than you expect, and it
  won't be obvious that's what happened.

## Next

<CardGroup cols={2}>
  <Card title="Filtering jobs" icon="Funnel" href="/docs/guides/filtering">
    The full filter vocabulary and every matching mode.
  </Card>
  <Card title="Enums & vocabularies" icon="ListTree" href="/docs/reference/enums">
    Every exact-match value, including the taxonomy list.
  </Card>
</CardGroup>
