--- name: huggingface-models-scraper description: List Hugging Face Hub repositories as structured records via the Apify Actor arman-bd/huggingface-models-scraper. Returns the repo id, author, hub URL, 30-day and all-time downloads, likes, trending score, pipeline task, library, space SDK, licence, the full tag list, gating status and both repository timestamps, for models, datasets or spaces. Use when a task needs model or dataset discovery, adoption and trending tracking, release watching for an organisation, or a licence and gating audit across many repos. Not for model card text, repository files, weights, inference or anything private. --- # Hugging Face Scraper: Models, Datasets & Spaces Apify Actor `arman-bd/huggingface-models-scraper`. Give it search terms, or no terms at all, and it walks the Hub's listing index in sort order and returns one dataset record per repository. Models, datasets and spaces come from the same input, chosen with `resourceType`. You supply no credentials. ## When to use it - Model or dataset discovery: the top N for a term, a task, or an organisation name. - Adoption tracking on a schedule, diffing `downloadsAllTime` and `likes` over time. - Release watching: sort by `createdAt` and search a lab's name to see what just shipped. - Licence and gating audits over a candidate list before anything gets shipped. - Trend reporting: sort by `trendingScore` for what the community picked up this week. ## When not to use it - Model card text, README content, config files, weights or any repository file. The listing index carries metadata only. - Inference, evaluation or benchmark results. - Private or organisation-internal repos. There is no token input. - Exact-id lookups of a handful of known repos. Search is a listing filter, not a fetch-by-id, and a term can return neighbours as well as the repo you meant. ## Call it ```js import { ApifyClient } from 'apify-client'; const client = new ApifyClient({ token: process.env.APIFY_TOKEN }); const run = await client.actor('arman-bd/huggingface-models-scraper').call({ searchQueries: ['llama'], resourceType: 'models', pipelineTag: 'text-generation', sortBy: 'downloads', maxResults: 100, }); const { items } = await client.dataset(run.defaultDatasetId).listItems(); const { value: summary } = await client .keyValueStore(run.defaultKeyValueStoreId) .getRecord('RUN_SUMMARY'); ``` One-shot over HTTP, when you want the rows back in the same request: ```bash curl -X POST "https://api.apify.com/v2/acts/arman-bd~huggingface-models-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \ -H "Content-Type: application/json" \ -d '{"searchQueries":["llama"],"pipelineTag":"text-generation","sortBy":"downloads","maxResults":50}' ``` The Actor is also exposed through Apify's MCP server as `arman-bd/huggingface-models-scraper`, so an MCP-capable agent can call it with no extra wiring. ## Input Every field is optional. With no input at all you get the 100 most downloaded models. | Field | Type | Required | Default | Notes | |---|---|---|---|---| | `searchQueries` | string[] | no | `[]` | Free-text terms matched against repository names and metadata. Each term is its own pass, so two terms return two sets. Duplicates and blanks are removed. Empty means walk the listing in sort order with no term filter. | | `resourceType` | string | no | `"models"` | One of `models`, `datasets`, `spaces`. Decides which index is read and which fields carry data. An unrecognised value throws before any request. | | `pipelineTag` | string | no | `""` | Task filter, e.g. `text-generation`, `automatic-speech-recognition`, `image-classification`. Models only: it is dropped with a warning for datasets and spaces. Empty means every task. | | `sortBy` | string | no | `"downloads"` | One of `downloads`, `likes`, `lastModified`, `createdAt`, `trendingScore`. Always descending. `downloads` on `spaces` silently becomes `likes`. An unrecognised value throws. | | `maxResults` | integer | no | `100` | Cap on records saved **per search query**, minimum `1`. With no queries it caps the single listing pass. Records arrive 100 per request, so 500 costs five requests. | **`sortBy` is doing more work than it looks like it is.** `maxResults` is a top-N cut, so the sort is the actual selector: the same term with `downloads` and with `createdAt` returns two disjoint sets of repos, not the same set reordered. Decide what "top" means before raising the cap. Cost is linear and cheap either way, one request per 100 records per query, so the real budget question is how many rows you want to store and process, not how many requests you can afford. ## Output One record per repository, per query pass. | Field | Type | Notes | |---|---|---| | `id` | string | Repository id, e.g. `meta-llama/Llama-3.2-1B-Instruct`. The join key. | | `resourceType` | string | `models`, `datasets` or `spaces`, echoed onto every row. | | `author` | string \| null | Owning user or organisation. Derived from `id` when the source omits it. | | `url` | string | Hub page for the repository, built for the right resource type. | | `downloads` | number \| null | Downloads in the **last 30 days**, not lifetime. `null` for spaces. | | `downloadsAllTime` | number \| null | Downloads since creation. `null` for spaces. | | `likes` | number \| null | Hub likes. The only engagement counter spaces have. | | `trendingScore` | number \| null | Current trending score. Short-lived and not comparable across weeks. | | `pipelineTag` | string \| null | Task, e.g. `text-generation`. Models only, `null` otherwise. | | `libraryName` | string \| null | Library, e.g. `transformers`. Models only, `null` otherwise. | | `sdk` | string \| null | Space runtime, e.g. `gradio`, `streamlit`, `docker`. Spaces only, `null` otherwise. | | `license` | string \| null | Licence identifier lifted out of the `license:` tag. `null` when the repo declares none. Not always an SPDX id: bespoke licences appear verbatim, e.g. `llama3.2`. | | `tags` | string[] | Full tag list: languages, datasets, paper ids, base models, regions, and the `license:` tag the licence was taken from. | | `gated` | boolean | Whether access is restricted. Gated repos are kept, with full public metadata. | | `gatedType` | string \| null | `auto` or `manual` when the source names an approval mode, otherwise `null`, including for some gated repos. | | `private` | boolean | Visibility flag. Always `false` in practice, since only public repos are listed. | | `createdAt` | string \| null | Repository creation time, ISO 8601. | | `lastModified` | string \| null | Last repository change, ISO 8601. The maintenance signal. | | `scrapedAt` | string | Run timestamp, ISO 8601. | A real record, tags trimmed: ```json { "id": "meta-llama/Llama-3.2-1B-Instruct", "resourceType": "models", "author": "meta-llama", "url": "https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct", "downloads": 10016506, "downloadsAllTime": 93318314, "likes": 1554, "trendingScore": 2, "pipelineTag": "text-generation", "libraryName": "transformers", "sdk": null, "license": "llama3.2", "tags": ["transformers", "safetensors", "llama", "text-generation", "conversational", "license:llama3.2", "…"], "gated": true, "gatedType": "manual", "private": false, "createdAt": "2024-09-18T15:12:47.000Z", "lastModified": "2024-10-24T15:07:51.000Z", "scrapedAt": "2026-08-06T11:35:15.076Z" } ``` ## RUN_SUMMARY Written to the run's key-value store under the key `RUN_SUMMARY`. **Read it.** It is the only place the Actor tells you that it overrode one of your inputs. ```json { "resourceType": "models", "queriesRequested": 2, "queriesFailed": 0, "failures": [], "itemsSaved": 150, "filters": { "searchQueries": ["whisper", "wav2vec"], "pipelineTag": "automatic-speech-recognition", "sortBy": "likes", "maxResults": 100 }, "finishedAt": "2026-08-06T11:35:22.410Z" } ``` Compare `filters` against what you sent: `sortBy` and `pipelineTag` are the effective values after the fallbacks below, so a mismatch there explains a surprising result set. `itemsSaved` well under `queriesRequested × maxResults` simply means the index ran out of matches, which is normal for a narrow term. `queriesFailed` above zero means whole query passes produced nothing; the run throws only when every pass failed. With no search terms, `queriesRequested` is `1` and any failure is labelled `all models`, `all datasets` or `all spaces`. ## Behaviour to plan around - **`maxResults` is per query, not per run.** Three terms at 200 is up to 600 records. Size your storage and your downstream loop on the product. - **Nothing is de-duplicated across queries.** Two terms that both match a repo produce two identical dataset items with different `scrapedAt` values. De-duplicate on `id` yourself when you pass more than one term. - **`downloads` is a rolling 30-day window.** Diffing it between two runs measures the change in recent popularity, not new downloads. `downloadsAllTime` is the cumulative counter, and is the one to difference over time. - **Sorting spaces by downloads silently becomes sorting by likes.** Spaces have no download counter. The run does not fail, and `filters.sortBy` in the summary is the only record of the substitution. - **`pipelineTag` is dropped for datasets and spaces**, again silently as far as the dataset is concerned. It comes back as `null` in `filters`, so a "task filter did nothing" result is explained there. - **Field availability follows `resourceType`.** Datasets have no `pipelineTag` or `libraryName`; spaces have no `downloads` or `downloadsAllTime` but do have `sdk`. These arrive as `null`, not as missing keys, so the columns stay stable in CSV. - **`license` is derived, not native.** It is parsed out of the tag list, so `null` means the repo carries no `license:` tag rather than that it is unlicensed, and the value can be a project-specific string rather than an SPDX id. For an audit, treat `null` as "unknown, go and look". - **Gated repos are included, not skipped.** The metadata is public even when the weights are not, so counts stay honest. Filter on `gated` if you need shippable-only. - **`trendingScore` is a moving target.** It reflects the current window and is not comparable between runs a week apart. Rank within one run only. - **Transient failures are retried** three times with linear backoff. A rejected query is final for that pass and is not retried; the remaining passes continue. ## Recipes **Top models for a term, ready to compare.** ```json { "searchQueries": ["llama"], "resourceType": "models", "sortBy": "downloads", "maxResults": 100 } ``` Rank on `downloads`, then filter out rows where `gated` is true or `license` is null if the shortlist has to be deployable. **Speech models by community signal, two terms.** ```json { "searchQueries": ["whisper", "wav2vec"], "resourceType": "models", "pipelineTag": "automatic-speech-recognition", "sortBy": "likes", "maxResults": 50 } ``` De-duplicate the combined 100 rows on `id` before ranking, since both terms can match the same repo. **Recently updated datasets, no search term.** The single listing pass walks the index in sort order. ```json { "resourceType": "datasets", "sortBy": "lastModified", "maxResults": 200 } ``` Use `lastModified` against your last run's timestamp to pick up only what moved. **Licence and gating audit for a shortlist.** Search the organisation, keep the columns that decide shipping. ```json { "searchQueries": ["mistralai"], "resourceType": "models", "sortBy": "downloads", "maxResults": 100 } ``` Keep `id`, `license`, `gated`, `gatedType` and `lastModified`; anything with `license` null needs a manual check rather than a default assumption.