Skip to content

DiscoGen API

DiscoGen is an AI research agent that runs your prompt against rich company or contact data at scale. Each record gets its own isolated context window (company profile, homepage text, or contact details), so the model sees focused, accurate information rather than a polluted batch. Optionally enable web search for live data enrichment.

Requires an LLM provider integration. Optionally pair with a search provider integration for BYOS (Bring Your Own Search) web search.

POST https://api.discolike.com/v1/discogen/process
POST https://api.discolike.com/v1/discogen/process-personas
GET https://api.discolike.com/v1/discogen/models
GET https://api.discolike.com/v1/discogen/status/{task_id}
DELETE https://api.discolike.com/v1/discogen/cancel/{task_id}

Both /process and /process-personas share these parameters:

ParameterTypeRequiredDescription
queryStringYesThe prompt to run against each record. Must be non-empty
integration_idStringNoLLM provider integration UUID. Uses your default LLM provider if omitted. Must be a valid UUID, or native-icp on /process to score an ICP validation prompt with DiscoLike’s own model (see Native ICP-Fit Scoring)
web_searchBooleanNoEnable web search enrichment (default: false). Returns 400 if no search provider is available and the model lacks built-in web search
search_provider_idStringNoSearch provider integration UUID for BYOS web search. Uses your default search provider if omitted. Must be a valid UUID
search_context_sizeStringNoNumber of search queries per record: low (default), medium, or high
include_x_searchBooleanNoInclude X (Twitter) search results for xAI models (default: false)
typed_columnsBooleanNoAnswer yes/no, fixed-set, and scale parts of your query with a calibrated judgment model instead of generated text (default: false). See Typed Columns
include_confidenceBooleanNoAdd a confidence column beside each typed column (default: false). No effect without typed_columns: true

Run an LLM prompt against each domain with company context.

ParameterTypeRequiredDescription
domainsArrayYesDomains to process (1-10,000)
context_modeStringNoWhat data the model sees per domain (default: website)
previous_discogen_dataObjectNoPrevious DiscoGen results keyed by domain, appended to context for iterative enrichment

Controls what data is included in each domain’s context window. Richer context produces better results but costs more. Credits are only charged for new domains; any domain you’ve already discovered in the last 90 days is free.

ModeContext IncludedCost
websiteCompany profile (firmographics, keywords, categories) + homepage textCredits (new domains only) + model fees
profileCompany profile only (firmographics, keywords, categories, summary)Credits (new domains only) + model fees
domainDomain name only; model relies on its own knowledge and web search if enabledModel fees only
Terminal window
curl -X POST "https://api.discolike.com/v1/discogen/process" \
-H "x-discolike-key: API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "What products or services does this company offer?",
"domains": ["acmecorp.com", "globex.com"],
"context_mode": "website"
}'
{
"task_id": "f47ac10b-58cc-4372-a567-0e02b2c3d479",
"column_name": "Products and Services",
"status": "in_progress",
"total_domains": 2
}

/process also accepts integration_id: "native-icp", which scores each domain with DiscoLike’s own ICP-fit model instead of your LLM provider. Your query must be an ICP validation prompt listing its criteria under Mandatory:, Reject if: and Nice-to-have:; the model answers nothing else. Validate generates such a prompt for you from plain ICP text, and is the easier entry point.

The run needs no LLM provider integration and makes no calls on your provider key. web_search is ignored, context_mode is always website, and column_name is ["ICP Fit", "ICP Score", "Reasoning"] rather than a label detected from your prompt, with Reasoning carrying the confidence band rather than a rationale, because this engine scores rather than explains. Column values and errors are documented under Native ICP-Fit Model.

Run an LLM prompt against contact records. Same async task model as domain processing.

ParameterTypeRequiredDescription
persona_idsArrayYesPersona IDs to process (1-10,000)
context_modeStringNoWhat data the model sees per persona (default: profile)
previous_discogen_dataObjectNoPrevious results keyed by persona ID for iterative enrichment

Credits are only charged for new contacts; any contact you’ve already discovered in the last 90 days is free.

ModeContext IncludedCost
fullContact profile + employer company profile + company homepage textCredits (new contacts only) + model fees
companyContact profile + employer company profileCredits (new contacts only) + model fees
profileFull contact profile (title, department, seniority, skills)Credits (new contacts only) + model fees
profile_summaryCondensed contact profile, a shorter summary of the full profileCredits (new contacts only) + model fees
name_onlyContact name and LinkedIn URL only; model relies on its own knowledge and web searchModel fees only
Terminal window
curl -X POST "https://api.discolike.com/v1/discogen/process-personas" \
-H "x-discolike-key: API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "Is this person likely a decision-maker for software purchases?",
"persona_ids": [12345, 67890],
"context_mode": "full"
}'

By default every column in a DiscoGen result is generated prose from your LLM provider. typed_columns: true lets part of a query be answered as a calibrated judgment instead: a yes/no, one of a fixed set, or a position on an ordered scale. Those columns are answered by DiscoLike’s TypeSafe (Jev) judgment model rather than your LLM provider, and come back with a probability behind the scenes, not just a label.

There is no separate schema to define. You send typed_columns: true and DiscoGen’s detector reads your query text to decide which parts of it are judgments and which are prose — the query is the whole specification. For example:

"Does this company offer a free trial? What pricing model do they use (Freemium, Subscription, or One-time)? Describe their target customer in one sentence."

detects a yes/no column, a fixed-set column, and a free-text column in a single run, and answers the first two with TypeSafe while your LLM provider writes the third.

Requesting a typed column requires a TypeSafe integration on your account (see LLM Providers: TypeSafe), independent of whichever integration_id the rest of the query runs on. Without one, every typed cell in the result comes back as "Error: No TypeSafe integration is configured; add one to use typed columns" rather than failing the row or the run. Other failures render as "Error: <reason>" in the same way — a rejected key reads "Error: JevRequestError". Generated columns in the same run are unaffected.

Typed yes/no cells contain "Yes" or "No". Fixed-set cells contain the chosen option. Scale cells contain a numeric, probability-weighted position on the detected ordered scale: 0 is the first level, 1 the second, and fractional values fall between levels. The level descriptions are in response_format.typed_columns[].criteria.

Add include_confidence: true to get a <Column Name> confidence column beside each typed column, one of:

ValueMeaning
High confidenceStrong signal in the evidence
Medium confidenceModerate signal
Low confidence - the judgment is near its cutoff, treat it as unresolvedNear the decision boundary; treat as unresolved rather than trusting the label

Pass a TypeSafe integration’s UUID as integration_id to run the query typed-only: every column is answered by TypeSafe and no LLM provider is called at all, so the run makes no calls on an LLM key and needs no LLM provider integration. column_name becomes the typed columns’ labels, and anything the query would otherwise generate as prose is dropped rather than left unanswered.

  • integration_id must be picked explicitly. If a TypeSafe integration is resolved by falling back to your organization’s default, the request 400s with TYPESAFE_NOT_A_GENERATIVE_ENGINE — a TypeSafe integration cannot be set as the org default in the first place (see LLM Providers), but an existing default can still end up pointing at one if its provider is edited afterward, so this is enforced on every request.
  • If the query carries no judgment at all (nothing the detector reads as yes/no, fixed-set, or scale), the request 400s with TYPESAFE_NO_TYPED_COLUMNS — rephrase it as a judgment, or select an LLM integration instead.
Terminal window
curl -X POST "https://api.discolike.com/v1/discogen/process" \
-H "x-discolike-key: API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "Does this company offer a free trial? What pricing model do they use (Freemium, Subscription, or One-time)?",
"domains": ["acmecorp.com", "globex.com"],
"typed_columns": true,
"include_confidence": true
}'
{
"status": "completed",
"results": {
"acmecorp.com": {
"Free Trial": "Yes",
"Free Trial confidence": "High confidence",
"Pricing Model": "Subscription",
"Pricing Model confidence": "Medium confidence"
}
},
"response_format": {
"response_type": "free_text",
"field_labels": {},
"typed_columns": [
{
"key": "free_trial",
"label": "Free Trial",
"kind": "noul",
"instructions": "Whether the company offers a free trial of its product",
"criteria": {
"true": "The website advertises a free trial or trial period",
"false": "The website does not advertise a free trial or trial period"
}
},
{
"key": "pricing_model",
"label": "Pricing Model",
"kind": "choice",
"instructions": "The company's primary pricing model",
"criteria": {
"Freemium": "A free tier exists alongside paid plans",
"Subscription": "Access requires a recurring paid plan, no free tier",
"One-time": "A single upfront purchase, no recurring billing"
}
}
]
}
}

Judgment values are returned under their column labels in results (or interim_results while the task is running), alongside any generated fields. For a generated list of rows, the judgments repeat on every row for that company or contact. Cancelled runs retain the answers that finished.

response_format.typed_columns echoes the detected columns, including the criteria TypeSafe judged each one against, so you can see why a value came back the way it did. A column is never listed in both places: field_labels/json_schema cover only the generative columns (empty on a typed-only run), and typed_columns covers only the judgment columns.

Jev spend is not billed to any provider key you configure; it is folded into estimated_cost but is not broken out as its own line in cost_metadata, which only tracks your LLM provider and search provider calls.

Returns available LLM models grouped by provider, with web search support flags.

Terminal window
curl "https://api.discolike.com/v1/discogen/models" \
-H "x-discolike-key: API_KEY"
{
"models": {
"openai": [
{ "name": "gpt-5-mini", "supports_web_search": true },
{ "name": "gpt-5-nano", "supports_web_search": true },
{ "name": "gpt-4o", "supports_web_search": false },
{ "name": "gpt-4o-mini", "supports_web_search": false }
],
"anthropic": [
{ "name": "claude-sonnet-4-6", "supports_web_search": false }
]
}
}

Both process endpoints return a task that can be polled and cancelled.

Terminal window
curl "https://api.discolike.com/v1/discogen/status/{task_id}" \
-H "x-discolike-key: API_KEY"
{
"status": "in_progress",
"progress": 45,
"interim_results": {
"acmecorp.com": "Enterprise cloud security platform offering..."
},
"estimated_cost": 0.0032
}
{
"status": "completed",
"progress": 100,
"results": {
"acmecorp.com": "Enterprise cloud security platform offering threat detection, SIEM, and compliance tools.",
"globex.com": "Industrial supply chain management with logistics optimization and warehouse automation."
},
"estimated_cost": 0.0058,
"response_format": {
"response_type": "free_text",
"description": "Products and services description",
"field_labels": { "answer": "Products and Services" }
},
"warnings": [],
"cost_metadata": {
"openai/gpt-4o-mini": {
"calls": 2,
"search_calls": 0,
"prompt_tokens": 6140,
"completion_tokens": 212,
"est_cost_usd": 0.001048
},
"search_provider": {
"provider": "serper",
"search_model": "serper/search",
"queries_executed": 4,
"queries_succeeded": 4,
"est_cost_usd": 0.004
}
}
}

estimated_cost and cost_metadata describe spend on your own provider keys (LLM and search). They are best-effort estimates for display, not DiscoLike billing: providers may hide cached-token usage, and custom endpoints cannot be priced.

cost_metadata has one entry per provider/model the run called, plus a search_provider entry when a search provider was used:

KeyDescription
callsLLM calls made with this model
search_callsSearches performed by the model’s built-in search tool (OpenAI, xAI, Gemini grounding). Always 0 when a search provider is configured, because the provider supplies the search context and built-in search is turned off
prompt_tokens, completion_tokensToken usage reported by the provider; absent if the provider did not report it
est_cost_usdEstimated spend for that model, or for the search provider
search_provider.queries_executedSearch queries sent to your search provider (BYOS) across the whole run
search_provider.queries_succeededQueries that returned a response

Read search_provider.queries_executed, not search_calls, to confirm that web search ran on a BYOS run.

Completed contact generation runs also carry title_validation: "llm" when your LLM validated each contact’s title against icp_text, "none" when the native extractor ran and titles were not validated. The field is absent on regular DiscoGen and persona runs. See Discover Contacts: BYOK Generate.

warnings lists non-fatal problems that changed how the run was grounded: the search provider running out of credits or tripping its circuit breaker, or web_search: true with a search provider configured that never executed a query. Results still complete, but affected rows are grounded only in profile and website text. The list is deduplicated and keeps the most recent entries.

Terminal window
curl -X DELETE "https://api.discolike.com/v1/discogen/cancel/{task_id}" \
-H "x-discolike-key: API_KEY"
{
"success": true,
"message": "Task cancellation requested",
"task_id": "f47ac10b-58cc-4372-a567-0e02b2c3d479",
"status": "cancelling"
}