This is what a creator database actually contains, how to evaluate one, and the failure modes that make a big number on a homepage meaningless.
The four things called a creator database
A profile index. A large crawl of public profiles, filterable. Coverage is broad, freshness varies, and nobody has verified that any given creator wants brand work. Most products marketed as influencer databases are this.
A curated roster. A smaller set, vetted or opted in. Coverage is narrower, intent is confirmed, and quality signal is higher per record.
A marketplace's supply. The creators available inside a platform you order through. You do not query it so much as order against it, and you cannot take the relationship with you.
Your own CRM. Creators you have briefed, paid and rated. The smallest and by far the most valuable, because it carries outcomes: who delivered on time, whose footage converted, who you would rebook.
Mature programmes run a funnel across these: index or roster for discovery, then promote the ones that work into a CRM you own. If a vendor only sells you the first and has no way to record outcomes, you are starting from zero every quarter.
What a record should contain
Follower count is the least useful field in the row, and it is the one every homepage leads with. A record worth querying carries:
| Field | Why it decides anything |
|---|---|
| Stable id | Lets you store the creator and re-read them later without re-resolving a handle that can change |
| Handle and name | Obvious, but the handle is not a stable key |
| Follower count | A band filter, not a quality signal |
| Engagement rate | The actual quality proxy, and only comparable if the database computes it consistently |
| Niche / category | Must be more reliable than a hashtag, which is a noisy proxy |
| Location (country, city, state) | Accent, seasonality, product availability, shipping |
| Gender, age | Only where you genuinely need it for casting |
| Contact email | The field that turns a list into a shortlist |
| Bio and external URL | Context, plus the link-in-bio that often reveals how they work |
| Verified flag | Weak signal, occasionally useful |
Engagement rate is a methodology, not a fact
No platform publishes an engagement rate. Every database computes one, and the choices are consequential: how many recent posts to average, mean or median, whether Reels are weighted against static posts, whether views or followers sit in the denominator, how outliers are handled.
Two databases can report different engagement for the same creator and both be defensible. The practical implications:
- Do not compare engagement figures across vendors. Compare the ordering they produce.
- Prefer a precomputed, consistent figure over computing your own, unless you will apply the same method forever. Your own numbers drift as your code changes, which makes historical comparison useless.
- Be suspicious of suspiciously high averages. If a database's median creator shows 8% engagement, something in the method is flattering.
Contact data is where lists die
A shortlist of 500 perfect creators with no email is a research project. Be precise about what a vendor means by contact data, because "includes contact info" covers three very different things:
- The platform's own public field. Instagram exposes a business email only when the creator fills it in on a professional account, and it is frequently empty. TikTok exposes no email at all. Pass-through of this is the most common implementation and the weakest.
- Bio parsing. Extracting an address a creator typed into their bio. Works, incomplete, decays.
- Resolved and stored contact, queryable as a filter.
Only the third lets you ask for creators you can reach. The test is simple: can you filter on it? On the Roster API that is has_email=true, one of 14 filters on GET /v1/creators/search:
If contact is a column you only discover after exporting, it is not a filter and your usable yield is whatever the coverage happens to be.
Why the headline number tells you almost nothing
Database sizes in this category run from a few hundred thousand to several hundred million profiles, and they are self-reported. A bigger index is not a better one, for four reasons:
Most of it is not addressable. A 200-million-profile index counts every public account, the overwhelming majority of which are private individuals who will never do brand work. The number you care about is how many match your brief and are contactable, which is a tiny fraction and never advertised.
Coverage is uneven by market and niche. A database can be enormous and still thin on, say, Spanish-language home-fitness creators in the 20,000 to 80,000 band. Size is a global claim; your brief is a local one.
Freshness decays. Follower counts, engagement and bios move constantly. An index that is broad but refreshed slowly returns confidently wrong numbers, which is worse than returning nothing.
Platform mix matters more than total. Instagram plus TikTok plus YouTube plus X summed into one figure hides whether the platform you actually advertise on is well covered.
Test it instead. Take a brief you have already run by hand and express it as filters. Count how many results are genuinely on-brief and contactable. That number, divided by the plan price, is the only size metric that means anything.
How to evaluate one in an afternoon
- Run one real brief, not a smoke test. Use a sourcing brief you ran manually and know the right answer to.
- Check a known cohort. Look up 200 creators you already have opinions about and measure coverage and accuracy. Every index has gaps; find yours before you depend on it.
- Measure contactable yield. Of the on-brief results, how many have a usable email? This is the number that predicts your sourcing cost.
- Spot-check freshness. Compare a sample of follower counts against the live profiles. Note how stale they are.
- Compare engagement ordering, not values, against your own judgement of those creators.
- Price the whole feature. Include whatever you have to build on top: normalization, engagement computation, contact enrichment, refresh scheduling.
Where databases stop
Every product in this category hands you a list and stops. The work after the list is briefing, contracting, shipping product, chasing deliverables, reviewing assets, tracking what published, and paying people.
That is fine at five creators and a real system at five hundred. If you are building it rather than staffing it, the UGC Roster API exposes creator search, briefs, campaigns and payouts behind one key at UGCRoster, which is why the endpoint list continues past search: /v1/roster, /v1/briefs, /v1/campaigns, /v1/contracts, /v1/shipments, /v1/deliverables, /v1/content, /v1/payouts, /v1/webhooks.
Credits are the only meter: a profile read is 1, a filtered search returning up to 100 records is 10, on published plans of $98 for 25,000 credits, $398 for 150,000 and $998 for 500,000.
More on the developer path in the influencer database API guide, on sourcing channels in where to find UGC creators, and on the raw-data alternatives in how creator data APIs compare.
FAQ
What is a UGC creator database?
A searchable index of creator profiles you can filter by attributes such as follower range, engagement rate, niche, location and whether a contact email is on file. The term covers four different products though: a broad scraped profile index, a smaller curated or opted-in roster, a marketplace's internal supply, and your own CRM of creators you have already worked with. They differ in coverage, freshness and whether creator intent has been confirmed.
How big should a creator database be?
Size is close to meaningless on its own, and it is always self-reported. Large indexes count every public account, most of which are private individuals who will never do brand work. What matters is how many records match your specific brief and carry a usable contact, which is a small fraction of any headline number. Test with a real brief and measure contactable yield rather than comparing homepage figures.
Do creator databases include email addresses?
Some do, many only pass through the platform's own public field. Instagram exposes a business email only when a creator has filled it in on a professional account, and it is frequently empty; TikTok exposes none. The meaningful test is whether contact availability is a filter you can query rather than a column you discover after exporting, because a list of creators you cannot reach is not a shortlist.
Why do two databases show different engagement rates for the same creator?
Because no platform publishes an engagement rate, so every database computes one, and the method involves real choices: how many recent posts to average, mean or median, how Reels are weighted against static posts, and whether followers or views are the denominator. Both figures can be defensible. Compare the ordering two databases produce rather than the numbers themselves.
Is a creator database better than a UGC marketplace?
They answer different questions. A marketplace is faster for low volume because matching, payment and delivery are handled, but the platform owns the relationship and your creative variety is capped by their supply. A database is cheaper per usable creator at volume and you keep the relationship, at the cost of running briefing, contracting and payment yourself.
How do I keep creator data fresh?
Decide a staleness policy per field and re-read on a schedule, because follower counts and engagement move constantly while niche and location rarely do. Store a stable id rather than a handle so a renamed account is still resolvable, record the timestamp on every read, and refresh the creators you are actively considering more often than the long tail you are not.