- Brand
- DataHub
- Category
- Data & Analytics
- Primary Subcategory
- Enterprise Knowledge Search & AI Context Layer
Integration details
Description
DataHub's MCP server gives AI agents the enterprise context they need to work with your data. Surface curated knowledge — runbooks, FAQs, business definitions, and vocabularies — so agents operate with the same shared understanding as your teams. Search across datasets, dashboards, and pipelines, then pull ownership, governance policies, quality signals, and documentation to understand what you're looking at. Trace lineage at the table and column level. Surface real SQL queries to see how data is actually used. Apply tags, glossary terms, owners, and descriptions at scale. The context layer that makes AI agents enterprise-ready.
- Integration type
- Plugin
- Verification status
- Not applicable
- Platform
- ChatGPT
- Primary Subcategory
- Enterprise Knowledge Search & AI Context Layer
- Secondary Subcategories
- None listed
- Brand
- DataHub
- Access
- Account required
- First tracked
- 2026-10-10
- Tool count
- 50
- Geography
- US
The Primary Subcategory used for this profile’s headline score.
Other Subcategories where the Integration is listed.
Your score is coming
ChatGPT now suggests Plugins on its own when they match a user's request.Your Plugin Discovery Score measures how often yours appears, and it will show here as soon as it’s ready.
What discovery looks like

Get alerts for DataHub Cloud
Get updates when DataHub Cloud’s Discoverability Score or category rank changes.
Competing in ChatGPT Enterprise Knowledge Search & AI Context Layer
View Category50 tools agents can invoke
Accept or reject one or more governance proposals (ActionRequests). decision: "accept" applies the proposed change; "reject" closes the proposal without change. note: optional explanation recorded on the ActionRequest.
accept_or_reject_proposals
Add one or more owners to multiple DataHub entities. This tool allows you to assign multiple entities with multiple owners in a single operation. Useful for bulk ownership assignment operations like assigning data stewards, technical owners, or business owners to datasets, dashboards, and other DataHub entities. Note: Ownership in DataHub is entity-level only. For field-level metadata, use tags or glossary terms instead. Args: owner_urns: List of owner URNs to add (must be CorpUser or CorpGroup URNs). Examples: ["urn:li:corpuser:john.doe", "urn:li:corpGroup:data-engineering"] entity_urns: List of entity URNs to assign ownership to (e.g., dataset URNs, dashboard URNs) ownership_type: The type of ownership to assign. Accepts either: - A built-in OwnershipType enum value: - TECHNICAL_OWNER: Involved in production, maintenance, or distribution - BUSINESS_OWNER: Principle stakeholders or domain experts - DATA_STEWARD: Involved in governance - A custom ownership type, passed by its user-facing name (e.g. "Producer"). Custom types are looked up by name (case-insensitive) via GraphQL and must already exist. Returns: Dictionary with: - success: Boolean indicating if the operation succeeded - message: Success or error message Examples: # Add technical owners to multiple datasets add_owners( owner_urns=["urn:li:corpuser:john.doe", "urn:li:corpGroup:data-engineering"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.customers,PROD)" ], ownership_type=TECHNICAL_OWNER ) # Add business owner add_owners( owner_urns=["urn:li:corpuser:jane.smith"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.customers,PROD)" ], ownership_type=BUSINESS_OWNER ) # Add data steward to multiple entities add_owners( owner_urns=["urn:li:corpuser:data.steward"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.sales,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.transactions,PROD)", "urn:li:dashboard:(urn:li:dataPlatform:looker,sales_dashboard,PROD)" ], ownership_type=DATA_STEWARD )
add_owners
Add semantic relationship(s) from a GlossaryTerm to one or more other GlossaryTerms. Use this to establish governed semantic relationships between terms: synonyms (aliases that drive search), antonyms (conflicting definitions), translations (multi-language equivalents), valid values (controlled vocabulary), and other structural relationships. Relationship types and their semantics (use the camelCase values exactly as shown): - "synonymOf" — terms that mean the same thing (drives cross-term search resolution) - "antonymOf" — terms that conflict in business meaning (triggers disambiguation in Ask DataHub) - "translatesTo" — term_urn is the canonical term; related_term_urns are translations in other languages - "hasValue" — term_urn is a dimension/concept; related_term_urns are its valid column values (e.g. Country hasValue CH, DE, FR — enables SQL value grounding) - "isRelatedTo" — general bidirectional relationship - "isA" — inheritance / "is a subtype of" - "hasA" — composition / "contains" Args: term_urn: URN of the source GlossaryTerm (e.g., "urn:li:glossaryTerm:Revenue"). related_term_urns: List of target GlossaryTerm URNs to link to. relationship_type: One of "synonymOf", "antonymOf", "translatesTo", "hasValue", "isRelatedTo", "isA", "hasA". Returns: Dictionary with: - success: True if the relationship(s) were added - term_urn: The source term URN - related_term_urns: The target term URNs - relationship_type: The relationship type used Examples: # Add a German translation add_related_terms( term_urn="urn:li:glossaryTerm:MonthlyRecurringRevenue", related_term_urns=["urn:li:glossaryTerm:MonatlichWiederkehrenderUmsatz"], relationship_type="translatesTo" ) # Mark two terms as synonyms add_related_terms( term_urn="urn:li:glossaryTerm:Revenue", related_term_urns=["urn:li:glossaryTerm:MRR", "urn:li:glossaryTerm:MonthlyRevenue"], relationship_type="synonymOf" ) # Add valid values for a dimension term add_related_terms( term_urn="urn:li:glossaryTerm:Country", related_term_urns=["urn:li:glossaryTerm:CH", "urn:li:glossaryTerm:DE"], relationship_type="hasValue" )
add_related_terms
Add structured properties with values to multiple DataHub entities. This tool allows you to assign structured properties to multiple entities in a single operation. Structured properties are schema-defined metadata fields that can store typed values (strings, numbers, etc.). Args: property_values: Dictionary mapping structured property URNs to lists of values. Example: { "urn:li:structuredProperty:io.acryl.privacy.retentionTime": ["90"], "urn:li:structuredProperty:io.acryl.common.businessCriticality": ["HIGH"] } entity_urns: List of entity URNs to assign properties to (e.g., dataset URNs, dashboard URNs) Examples: # Add retention time and criticality to datasets add_structured_properties( property_values={ "urn:li:structuredProperty:io.acryl.privacy.retentionTime": ["90"], "urn:li:structuredProperty:io.acryl.common.businessCriticality": ["HIGH"] }, entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.customers,PROD)" ] ) # Add numeric property add_structured_properties( property_values={ "urn:li:structuredProperty:io.acryl.dataQuality.scoreThreshold": [0.95] }, entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.verified_data,PROD)" ] ) # Add multiple values for a multi-valued property add_structured_properties( property_values={ "urn:li:structuredProperty:io.acryl.common.dataClassification": ["PII", "SENSITIVE"] }, entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)" ] )
add_structured_properties
Add one or more tags to multiple DataHub entities or their columns (e.g., schema fields). This tool allows you to tag multiple entities or their columns with multiple tags in a single operation. Useful for bulk tagging operations like marking multiple datasets as PII, deprecated, or applying governance classifications. Args: tag_urns: List of tag URNs to add (e.g., ["urn:li:tag:PII", "urn:li:tag:Sensitive"]) entity_urns: List of entity URNs to tag (e.g., dataset URNs, dashboard URNs) column_paths: Optional list of column_path identifiers (e.g., column names for schema fields). Must be same length as entity_urns if provided. Use None or empty string for entity-level tags. For column-level tags, provide the column name (e.g., "email_address"). Verify that the column_paths are correct and valid via the schemaMetadata. Use get_entity tool to verify. Returns: Dictionary with: - success: Boolean indicating if the operation succeeded - message: Success or error message Examples: # Add tags to multiple datasets add_tags( tag_urns=["urn:li:tag:PII", "urn:li:tag:Sensitive"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.customers,PROD)" ] ) # Add tags to specific columns add_tags( tag_urns=["urn:li:tag:PII"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)" ], column_paths=["email", "phone_number"] ) # Mix entity-level and column-level tags add_tags( tag_urns=["urn:li:tag:Deprecated"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.old_table,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)" ], column_paths=[None, "deprecated_column"] # Tag whole table and a specific column )
add_tags
Add one or more glossary terms (terms) to multiple DataHub entities or their columns (e.g., schema fields). This tool allows you to associate multiple entities or their columns with multiple glossary terms in a single operation. Useful for bulk term assignment operations like applying business definitions, standardizing terminology, or enriching metadata with domain knowledge. Args: term_urns: List of glossary term URNs to add (e.g., ["urn:li:glossaryTerm:CustomerData", "urn:li:glossaryTerm:SensitiveInfo"]) entity_urns: List of entity URNs to annotate (e.g., dataset URNs, dashboard URNs) column_paths: Optional list of column_path identifiers (e.g., column names for schema fields). Must be same length as entity_urns if provided. Use None or empty string for entity-level glossary terms. For column-level glossary terms, provide the column name (e.g., "customer_email"). Verify that the column_paths are correct and valid via the schemaMetadata. Use get_entity tool to verify. Returns: Dictionary with: - success: Boolean indicating if the operation succeeded - message: Success or error message Examples: # Add glossary terms to multiple datasets add_glossary_terms( term_urns=["urn:li:glossaryTerm:CustomerData", "urn:li:glossaryTerm:PersonalInformation"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.customers,PROD)" ] ) # Add glossary terms to specific columns add_glossary_terms( term_urns=["urn:li:glossaryTerm:EmailAddress"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)" ], column_paths=["email", "contact_email"] ) # Mix entity-level and column-level glossary terms add_glossary_terms( term_urns=["urn:li:glossaryTerm:Revenue"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.sales,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.transactions,PROD)" ], column_paths=[None, "total_amount"] # Term for whole table and a specific column )
add_terms
Ask the 'Ask DataHub' agent. DataHub's general data-catalog assistant — searches assets, explores lineage, inspects schemas, and answers questions about your data. This call is asynchronous: it returns a conversation_urn, not the agent's answer. Poll get_agent_response(conversation_urn=...) until the status is COMPLETED to read the response, and do not answer for the agent before then. Ask follow-ups with continue_agent_conversation.
ask_agent__ask_datahub
Ask another question in an existing agent conversation. The agent keeps everything already asked and answered, so use this for follow-ups rather than starting a new conversation. Like the first call this is asynchronous: poll get_agent_response with the same conversation_urn. Wait for the previous question to finish before asking another. Args: conversation_urn: The conversation to continue. query: The follow-up question for the agent. Returns: The conversation handle and its status. Poll it with get_agent_response.
continue_agent_conversation
Create an unpublished document that can be submitted for publication review. This creates the document but does not publish it. Call list_lifecycle_stages, then propose_lifecycle_stage with the user-assignable Published stage when the draft is ready for review. Do not use this tool to propose edits to an existing document; use propose_document_edit instead. The creator is added as an owner automatically. parent_document_urn controls where review routing begins for the initial publication proposal.
create_document_draft
Create a new (unversioned) GlossaryTerm in the Business Glossary. Use this to create a standalone term. If you want to create a versioned term from the start, use create_glossary_term_version instead. Args: name: Display name for the term (e.g., "Revenue"). description: Business definition / description for the term. parent_node_urn: Optional URN of the parent GlossaryNode (e.g., "urn:li:glossaryNode:KPIs"). Places the term under that node in the glossary hierarchy. id: Optional custom slug used as the URN key (e.g., "Revenue"). A UUID is generated if omitted. Avoid spaces; use underscores or CamelCase. Returns: Dictionary with: - success: True if creation succeeded - term_urn: URN of the newly created GlossaryTerm Examples: create_glossary_term( name="Revenue", description="Total revenue = sum of closed-won deals in the reporting period.", parent_node_urn="urn:li:glossaryNode:FinancialKPIs", id="Revenue" )
create_glossary_term
Draft SQL for a set of known tables using DataHub semantic understanding. PRECONDITION: The caller already knows which tables to query. Get URNs from: - ``find_sql_context`` returning ``suggested_tables`` when no anchor matched - ``smart_search`` / ``search_documents`` results - The user naming tables directly (e.g. "see revenue using orders and customers") If you need to figure out the query shape from analyst conventions instead, call ``find_sql_context`` first — that tool surfaces past-query patterns grounded in semantic anchor documents. HOW IT WORKS: Builds a semantic model from the given tables (schema, column classifications, relationships inferred from historical SYSTEM queries) and asks an LLM to synthesise SQL against it. Returns the SQL with confidence scores, explicit assumptions, ambiguities, and clarifying questions for the caller to surface. Args: natural_language_query: The user's question in natural language. Example: "Show me the top 10 customers by total order value last month" table_urns: List of dataset URNs to include in the semantic model. These should be the tables relevant to answering the query. Example: ["urn:li:dataset:(urn:li:dataPlatform:snowflake,prod.sales.orders,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,prod.sales.customers,PROD)"] platform: Target SQL dialect. Accepts any SQL platform DataHub supports (e.g. snowflake, bigquery, postgres, redshift, mysql, databricks, trino, presto, athena, oracle, mssql, clickhouse) or "generic" (default) for cross-dialect standard SQL. Dialect-specific prompt hints exist for snowflake/bigquery/postgres; other platforms fall back to generic hints, but the platform name is still given to the model so it can apply dialect-specific conventions. additional_context: Optional dict with additional context from the calling agent. Can include structured information like: - filters: {"date_range": "last 30 days", "region": "US"} - aggregation: "daily" or "monthly" - specific_columns: ["customer_id", "order_total"] - business_context: "We need this for the quarterly report" Returns: Dict with: - sql: The drafted SQL query - explanation: What the query does - platform: Target platform - confidence: "high", "medium", or "low" - tables_used: List of table URNs used - assumptions: List of assumptions made - ambiguities: List of ambiguities that could affect results - suggested_clarifications: Questions to ask for better results - semantic_model_summary: Stats about the semantic model used Examples: # Simple aggregation query draft_sql_for_tables( natural_language_query="What were total sales by region last quarter?", table_urns=["urn:li:dataset:(...):sales_orders"], platform="snowflake" ) # Multi-table query with context draft_sql_for_tables( natural_language_query="Which customers haven't ordered in 90 days?", table_urns=[ "urn:li:dataset:(...):customers", "urn:li:dataset:(...):orders" ], platform="bigquery", additional_context={"threshold_date": "2024-01-01"} )
draft_sql_for_tables
Surface a compact routing card for a SQL-flavoured question. Call this first for natural-language questions likely to require SQL. The response is retrieval evidence, not generated SQL: - ``authoritative_instructions``: search groups — each shows the ``search_documents`` call that already ran (verbatim) and the curated non-anchor documents it returned, every document as the excerpt that best matches the question (not the full body). Read these first; do not re-run those calls, and fetch a document's urn when its excerpt is not enough. - ``authoritative_instructions_note``: explains the completed searches and how to go beyond them. - ``candidate_tables``: table URNs with trust, measure, join, document, and anchor signals. - ``anchors``: compact trust-scaled anchor evidence; generated drafts carry facts only. - ``documents``: further documents grouped by the ``search_documents`` call that returned them, each as a query-matched excerpt plus link. Guards, date shapes, and value domains are served at hop 2 by ``inspect_tables_for_sql`` on the tables the query will actually use. ARGS: question: The user's complete natural-language question.
find_sql_context
Check on an agent conversation started by ask_agent__* and read its answer. Agent calls are asynchronous: the ask_agent__* tool returns a conversation_urn right away, and this tool reports where that conversation got to. Poll it until the status is no longer QUEUED or RUNNING. Args: conversation_urn: The conversation_urn returned by an ask_agent__* tool or by continue_agent_conversation. thinking_after: The thinking_cursor from your previous poll of this turn, so thinking holds only what is new. Use 0 on the first poll. Returns: status QUEUED or RUNNING - the agent is still working. thinking lists the agent's progress messages since thinking_after. Wait the number of seconds in poll_after_seconds, then call this tool again with thinking_after set to thinking_cursor. Do not answer for the agent while a turn is in progress. status COMPLETED - response holds the agent's answer. status FAILED or INTERRUPTED - error says what went wrong. The conversation history is intact, so you can ask again with continue_agent_conversation.
get_agent_response
Get SQL queries associated with a dataset or column to understand usage patterns. This tool retrieves actual SQL queries that reference a specific dataset or column. Useful for understanding how data is used, common JOIN patterns, typical filters, and aggregation logic. PARAMETERS: source - Filter by query origin: - "MANUAL": Queries written by users in query editors (real SQL patterns) - "SYSTEM": Queries extracted from BI tools/dashboards (production usage) - null: Return both types (default) COMMON USE CASES: 1. SQL Generation - Learn real query patterns: get_dataset_queries(urn, source="MANUAL", count=5-10) → See how users actually write SQL against this table → Discover common JOINs, aggregations, filters → Match organizational SQL conventions and patterns 2. Production usage analysis: get_dataset_queries(urn, source="SYSTEM", count=20) → See how dashboards and reports query this data → Understand which queries run in production → Identify critical query patterns 3. Column usage patterns: get_dataset_queries(urn, column="customer_id", source="MANUAL", count=5) → See how a specific column is used in queries → Learn filtering and grouping patterns for that column → Discover relationships via JOIN patterns 4. General usage exploration: get_dataset_queries(urn, count=10) → Get mix of manual and system queries → Understand overall table usage EXAMPLES: - Get manual queries for SQL generation: get_dataset_queries( urn="urn:li:dataset:(urn:li:dataPlatform:snowflake,prod.sales.orders,PROD)", source="MANUAL", count=10 ) - Get dashboard queries (production usage): get_dataset_queries( urn="urn:li:dataset:(...)", source="SYSTEM", count=20 ) - Column-specific query patterns: get_dataset_queries( urn="urn:li:dataset:(...)", column="created_at", source="MANUAL", count=5 ) RESPONSE STRUCTURE: - total: Total number of queries matching criteria - start: Starting offset - count: Number of results returned - queries: Array of query objects with: - urn: Query identifier - properties.statement.value: The actual SQL text - properties.statement.language: Query language (SQL, etc.) - properties.source: MANUAL or SYSTEM - properties.name: Optional query name - platform: Source platform - subjects: Referenced datasets/columns (deduplicated to dataset URNs) ANALYZING RETRIEVED QUERIES: Once you retrieve queries, examine the SQL statements to identify: - JOIN patterns: Which tables are joined? On what keys? - Aggregations: Common SUM, COUNT, AVG, GROUP BY patterns - Filters: Typical WHERE clauses, date range logic - Column usage: Which columns appear frequently vs rarely - CTEs and subqueries: Complex query structures BEST PRACTICES: - For SQL generation: Use source="MANUAL" (count=5-10) to see real user patterns - For production analysis: Use source="SYSTEM" to see dashboard/report queries - Start with moderate count (5-10) to avoid overwhelming context - If no queries found (total=0), proceed without query examples - not all tables have queries - Parse the SQL statements yourself to find patterns - they are not full-text searchable
get_dataset_queries
Get query-volume history for a dataset as UTC-dated points. Use this tool when the user asks how a dataset's query volume changes over time. Resolve the dataset URN first. The returned points can be summarized directly or passed unchanged to a host-supported line chart. Args: urn: Full dataset URN. time_range: DAY, WEEK, MONTH, or QUARTER. Defaults to MONTH. Returns: Dataset URN, requested range, total query count, and dated query-count points suitable for a line chart.
get_dataset_usage_history
Get detailed information about one or more entities by their DataHub URNs. IMPORTANT: Pass an array of URNs to retrieve multiple entities in a single call - this is much more efficient than calling this tool multiple times. When examining search results, always pass an array with the top 3-10 result URNs to compare and find the best match. Accepts an array of URNs or a single URN. Supports all entity types including datasets, assertions, incidents, dashboards, charts, users, groups, and more. The response fields vary based on the entity type. `fields` is an allowlist: identity always, anything else only when named — including the description, which can be long. What was withheld is listed in `omittedFields` with its size, so ask again naming it rather than assuming it is missing. For a wide table name the columns instead of the whole schema: schema_columns=["user_id"] (substring match). A whole-schema dump returns each column's description as a short preview; one ending in a `[+N chars]` marker is truncated — name that column in schema_columns to get its full definition. Descriptions written by propagation, by AI, or from Context docs live in `documentation`, not `description`. Asking for `description` returns both. Each `documentation` entry carries an `origin` — `propagated` and `ui_authored` are copies of text a human wrote, `document_extracted` is an excerpt of a human-written Context doc about this asset, `ai_generated` is model output that may be wrong.
get_entities
Get upstream or downstream lineage for any entity, including datasets, schemaFields, dashboards, charts, etc. Set upstream to True for upstream lineage, False for downstream lineage. Set `column: null` to get lineage for entire dataset or for entity type other than dataset. When column is set, results are always datasets (only datasets have schema fields). Setting max_hops to 3 is equivalent to unlimited hops. FILTER SYNTAX: SQL-like WHERE clause over the fields below. Supports AND, OR, NOT, parentheses, IN (a, b), comparison (>, >=, <, <=), and IS [NOT] NULL. e.g. entity_type = dataset AND env = PROD AND (platform = snowflake OR platform = bigquery) SUPPORTED FILTER FIELDS: - entity_type: dataset, dashboard, chart, corp_user, corp_group, dataProduct, etc. - entity_subtype (or subtype): Table, View, Model, etc. - platform: snowflake, bigquery, looker, tableau, etc. - domain, container, tag, glossary_term, owner: full URN (urn:li:...), found by searching entity_type = domain / container / tag / glossaryTerm / corp_user - env: PROD, DEV, STAGING (only use if explicitly requested) - status: NOT_SOFT_DELETED, ALL, ONLY_SOFT_DELETED - deprecated, hasActiveIncidents, hasFailingAssertions: true or false - columnCount: number of columns (from dataset profiling) - lastModifiedAt, createdAt: when the asset was last changed / created in the source system (datasets, dashboards, charts, pipelines, containers, etc.) - lastOperationTime: time of the last write seen in the source's audit logs (datasets only) Compare time fields to an ISO date, e.g. lastModifiedAt >= 2026-01-31 STRUCTURED PROPERTY FILTERS: use structuredProperties.<qualifiedName> as the field name, where qualifiedName is the ID portion of the structured property URN, e.g. structuredProperties.io.acryl.privacy.retentionTime = 30. Values are CASE-SENSITIVE; some properties have fixed allowed values (e.g. "Published"), so use get_entities on the property URN to check definition.allowedValues. Double-quote any value containing spaces, =, or parentheses: subtype = "Materialized View" subtype IN ("Materialized View", Table) PAGINATION: Pass the response's `nextOffset` back as `offset` to get the next page, and stop when `hasMore` is false. Do not step by max_results — the token budget can return fewer entities than requested, and stepping by max_results would skip the difference. QUERY PARAMETER - Search within lineage results: You can filter lineage results using the `query` parameter with same /q syntax as search tool: - /q my_schema -> find tables in specific schema - /q customer+transactions -> find entities with both terms - /q looker OR tableau -> find dashboards on either platform - /q * -> get all lineage results (default) Examples: - Find a specific table in a large downstream set: query="my_db.my_schema.events" - Find Looker dashboards in lineage: query="/q tag:looker" - Get all results: query="*" or omit parameter FIELDS - Control response size: Each node returns identity and description only. `fields` is an allowlist — name what the question needs, e.g. fields=["ownership"] for "who owns the downstream tables". Anything withheld is listed per node in `omittedFields`. Fewer fields per node also means more nodes fit within the token budget.
get_lineage
Get detailed lineage path(s) between two specific entities or columns. Returns the paths array from searchAcrossLineage, showing the exact transformation chain(s) including intermediate entities, columns, and transformation query URNs. Unlike get_lineage() which returns all lineage targets with compact lineageColumns, this tool focuses on ONE specific target and returns detailed path information. Args: source_urn: URN of the source dataset target_urn: URN of the target dataset source_column: Optional column name in source dataset target_column: Optional column name in target dataset (required if source_column provided) direction: Optional direction to search. If None (default), automatically discovers the path by trying downstream first, then upstream. Specify "downstream" or "upstream" explicitly for better performance if you know the direction. Returns: Dictionary with: - source: Source entity/column info - target: Target entity/column info - paths: Array of path objects from GraphQL (with QUERY URNs) - pathCount: Number of paths found Examples: # Column-level paths paths_result = get_lineage_paths_between( source_urn="urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.base_table,PROD)", target_urn="urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.final_table,PROD)", source_column="user_id", target_column="customer_id" ) # Returns paths with QUERY URNs showing transformation chain # Fetch SQL for specific query of interest query_details = get_entities(paths_result["paths"][0]["path"][1]["urn"]) # Dataset-level paths (auto-discover direction) get_lineage_paths_between( source_urn="urn:li:dataset:(...):base_table", target_urn="urn:li:dataset:(...):final_table" ) # Explicit direction for better performance get_lineage_paths_between( source_urn="urn:li:dataset:(...):base_table", target_urn="urn:li:dataset:(...):final_table", direction="downstream" )
get_lineage_paths_between
Get information about the currently authenticated user. This tool fetches detailed information about the authenticated user including: - User profile information (username, email, full name, etc.) - Platform privileges (what the user can do in DataHub) - Group memberships - User settings and preferences Returns: Dictionary with: - success: Boolean indicating if the operation succeeded - data: User information including corpUser and platformPrivileges - message: Success or error message Example: # Get current user information get_me()
get_me
Get the schema of DataHub's relationship graph: which relationship types exist between which entity types. Call this BEFORE traversing relationships to learn the vocabulary of this instance — including CUSTOM relationship types you cannot guess. Once per conversation is usually enough. WHAT YOU GET: A Turtle (RDF) ontology. Each rdf:Property subject is a relationship type — e.g. rel:DownstreamOf is urn:li:relationshipType:DownstreamOf — and its IRI is directly usable as a relationship_types value on the relationship traversal tools (built-in NATIVE types also accept the bare local name, e.g. "DownstreamOf"). Key predicates: - schema:domainIncludes / schema:rangeIncludes: source / destination entity types, as a UNION of independent declarations. - dh:path: one blank node per declaration with dh:source, dh:destination and per-declaration dh:isLineage / dh:isUpstream flags. Paths with dh:isLineage true are lineage edges — traversed by get_lineage and by include_lineage=true, and the right surface for data-flow/usage questions. - dh:origin: "NATIVE" (declared in the metadata model) or "CUSTOM" (authored on this instance). A CUSTOM property with no domain/range statements has unknown adjacency. - a owl:SymmetricProperty: traversed in both directions regardless of the requested direction. a owl:TransitiveProperty: chaining is meaningful. - dh:reverseDisplayName: display label for reverse traversal. SCOPING: Prefer entity_type (e.g. "dataset", "glossaryTerm") to restrict the export to relationship types whose domain or range includes that type — the full graph is a few hundred adjacency facts.
get_metadata_graph
Traverse DataHub's relationship graph from a starting entity. WHEN TO USE: Use this tool to explore NON-LINEAGE relationships, or a COMBINATION of non-lineage and lineage relationships in one walk (set include_lineage=true). For purely data-flow questions — upstream/downstream, "what uses or feeds this asset" — use get_lineage instead: usage is transitive, so those questions need a multi-hop lineage walk, and get_lineage's filter applies to final results only (unlike this tool's entity_types, which constrains every hop). RELATIONSHIP TYPES: relationship_types accepts built-in relationship names and urn:li:relationshipType:* / urn:li:structuredProperty:* urns, and ONLY the types you name are traversed (there is no "all types" wildcard). Do not guess type names: call get_metadata_graph first to see which relationship types exist between which entity types — including custom types defined on this instance — and choose from its results. DIRECTION SEMANTICS: Relationship edges point FROM the entity holding the metadata TO the entity it references. direction="incoming" from entity X returns entities whose edges point at X. The default "undirected" traverses both ways — use it when unsure of edge orientation, and pin incoming/outgoing for precision. LINEAGE IS A SEPARATE AXIS: relationship_types and direction apply ONLY to the named-relationship walk — they do not filter or constrain lineage edges in any way. When include_lineage=true, the full native lineage edge set is traversed as well, controlled solely by lineage_direction (omit for both upstream and downstream), and the two walks are unioned. relationship_types may be omitted only when include_lineage=true (a pure-lineage walk). ENTITY TYPES FILTER APPLIES AT EVERY HOP: entity_types constrains the entities visited at each hop, not just the final results — a multi-hop walk through an intermediate entity of a filtered-out type finds nothing beyond it. Prefer omitting entity_types when max_hops > 1. PAGINATION: Results are sorted by (degree, urn) and paged server-side. Pass the response's `nextOffset` back as `offset` to get the next page, and stop when `hasMore` is false. Do not step by max_results — the token budget can return fewer entities than requested, and stepping by max_results would skip the difference. FIELDS - Control response size: Each related entity returns identity and description only. `fields` is an allowlist — name what the question needs, e.g. fields=["ownership"] for "who owns the charts using this dataset". Anything withheld is listed per entity in `omittedFields`. Fewer fields per entity also means more entities fit within the token budget. If the response has `isPartial: true`, the server capped the walk (traversal limit or timeout) — narrow relationship_types/entity_types or reduce max_hops.
get_relationships
Search within document content using regex patterns. Similar to ripgrep/grep - finds matching excerpts within documents. Use search_documents() first to find relevant document URNs, then use this tool to search within their content. PATTERN SYNTAX (RE2 regex): - Simple text: "deploy" matches the word deploy - Case insensitive: "(?i)deploy" matches Deploy, DEPLOY, deploy - Word boundaries: r"\bword\b" matches whole word only - Alternatives: "deploy|release" matches either term - Wildcards: "deploy.*prod" matches deploy followed by prod - Character classes: "[Dd]eploy" matches Deploy or deploy PARAMETERS: urns: List of document URNs to search within - Get these from search_documents() results - Example: ["urn:li:document:doc1", "urn:li:document:doc2"] pattern: Regex pattern to search for - Examples: "kubernetes", "(?i)deploy.*production", "error|warning" - Use ".*" to get raw content (for continuing after truncation) - Runs on RE2: bounded repetition {m,n} is capped at 1000. Do NOT bake a large window into the pattern (e.g. "phrase[\s\S]{0,3000}") — use context_chars for surrounding text. Oversized repetitions are auto-clamped to 1000 with a note. context_chars: Characters to show before/after each match (default: 200) - Higher values show more surrounding context - When reading raw content (pattern=".*"), use higher values (e.g., 8000) - This is the right way to capture text around a match — not a {0,N} quantifier. max_matches_per_doc: Maximum excerpts to return per document (default: 5) - Limits output size for documents with many matches - A match inside an earlier excerpt is counted but not shown again, and excerpts never repeat text, so nearby matches share one excerpt start_offset: Character offset to start searching from (default: 0) - Use this to page through a long document: read one window, then call again with a higher start_offset to continue where you left off - When start_offset is set, the result includes content_length (the document's total size) so you know when the whole body has been read EXAMPLE WORKFLOWS: 1. Find deployment instructions: docs = search_documents(query="deployment", filter="platform = confluence") urns = [r["entity"]["urn"] for r in docs["searchResults"]] grep_documents(urns, pattern="kubectl apply", context_chars=300) 2. Find all error handling sections (case insensitive): grep_documents(urns, pattern="(?i)error|exception|failure") 3. Find specific configuration values: grep_documents(urns, pattern=r"timeout.*=.*\d+") 4. Read a long document's raw body in chunks: grep_documents(urns=[doc_urn], pattern=".*", context_chars=8000) # Returns consecutive excerpts from the start of the body, up to the output # limit. If the body is longer, the result's next_offset says where to continue: grep_documents(urns=[doc_urn], pattern=".*", context_chars=8000, start_offset=next_offset) # content_length reports the total OUTPUT LIMIT: One call returns at most about 60,000 characters of excerpt text, across all documents. Output stops at the limit, even mid-excerpt: the response then has truncated: true, and every document with more to show has a next_offset (documents not reached yet have matches: []). RETURNS: - results: List of documents with matches, each containing: - urn: Document URN - title: Document title - matches: List of excerpts with position info (positions are absolute). An excerpt cut by the output limit has truncated: true - total_matches: Total matches found (may exceed max_matches_per_doc) - content_length: Total length of document content (when start_offset is used, or the output limit cut this document) - next_offset: Present when this document has more to show (matches beyond max_matches_per_doc, or text past the output limit); pass it as start_offset to continue exactly there - total_matches: Total matches across all documents - documents_with_matches: Number of documents containing matches - truncated: Present (true) only when the output limit was reached. - truncation_note: Present with truncated; says how to continue. - skipped_urns: Present only when an input was not a usable URN. Each entry has `urn` (the rejected value) and `error` (why, and how to resolve it). The remaining URNs were still searched — resolve these with search_documents() and retry. - pattern_note: Present only when the pattern was rewritten to compile, explaining what changed. - error: Present instead of results when the pattern could not compile at all.
grep_documents
Search across all aspects of a DataHub entity using regex and return matching snippets. Use this as the FIRST step when exploring an entity — especially wide tables with hundreds of columns. Grep finds which columns, tags, properties, or descriptions mention your search terms, then use list_schema_fields or get_entities to get full details on the matches. Each string value in every aspect is searched individually — wildcards like .* stay inside one value and can't span across fields. Results are capped at 20 matches per aspect, 60K chars total. Match paths use identifier keys when available, e.g. schemaMetadata.fields[fieldPath=customer_email].description so you can see which field matched without a follow-up call. PATTERN SYNTAX (RE2 regex, case-sensitive by default): - Simple text: "deploy" matches the word deploy (case-sensitive) - Case-insensitive: "(?i)deploy" matches Deploy, DEPLOY, deploy - Word boundaries: r"\btenant\b" matches whole word only - Alternatives: "(?i)zip|postal|geography" matches any term - Wildcards: "(?i)tenant\w*id" matches tenant followed by id in one value WORKFLOW: grep_entities first to discover relevant fields, then list_schema_fields for full column details, or get_entities for structured extraction. Args: urn: Entity URN to search. pattern: RE2 regex pattern (case-sensitive by default, use (?i) for insensitive). context_chars: Characters of context around each match (default: 120, max: 2000). Returns: Dictionary with: - urn: The entity URN - pattern: The search pattern used - matches: Dict mapping aspect names to lists of {path, snippet} dicts - matchCount: Total regex hits across all aspects (may exceed snippet count) - aspectsSearched: Number of aspects that were searched - error: Present instead of matches when the pattern could not compile. Examples: # Find columns related to geography or tenants (case-insensitive) grep_entities( urn="urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.wide_table,PROD)", pattern="(?i)zip|tenant|geography" ) # Search for PII-related metadata across all aspects grep_entities( urn="urn:li:dataset:(urn:li:dataPlatform:bigquery,project.dataset.users,PROD)", pattern="(?i)email|phone|ssn|pii" )
grep_entities
Hydrate a table's conventions BEFORE writing SQL that uses it — always. Call this on EVERY table your query will SELECT FROM or JOIN, before writing the SQL, regardless of where you found the table. If you switch tables after reading a document redirect, inspect the newly chosen table too. For each table this returns corpus-observed guard filters, date shapes, value domains, and bounded body text from its curated registry document when one exists. Curated documentation outranks generated history. ARGS: table_urns: Full ``urn:li:dataset:...`` URNs of the tables the SQL will use. question: The user's natural-language question — always pass it. It drives the worked-example match; the SQL you write is only as good as the question here. Pass ALL the tables your query will use in a SINGLE call (not one at a time). When one saved query matches both the question and that set of tables, it is returned as `worked_example` — a construction template to adapt, not copy.
inspect_tables_for_sql
Return lifecycle stages evaluated for a target entity type. Call this BEFORE set_lifecycle_stage or propose_lifecycle_stage, passing the target GraphQL entity type (for example, "DOCUMENT" or "GLOSSARY_TERM"). Lifecycle stages are operator-defined and vary per deployment — never guess or hardcode a URN. Each item in the result contains: - urn: the value to pass as stage_urn to set_lifecycle_stage() - name: human-readable label (e.g. "Draft", "In Review", "Deprecated") - userAssignable: whether this stage can be set or proposed for entity_type - hideInSearch: true if entities in this stage are hidden from default search - allowedPreviousStages: URNs of stages that may transition into this one (null = any prior stage allowed; empty list = unreachable) - entityTypes: entity types this stage applies to (null = all types) Example: to move a glossary term to review, find the stage with name "In Review" and pass its urn to set_lifecycle_stage().
list_lifecycle_stages
List pending governance proposals with the content needed for review. proposal_type filters by ActionRequestType, e.g. "CREATE_GLOSSARY_TERM", "TERM_ASSOCIATION", "TAG_ASSOCIATION". Omit to return all pending proposals assigned to the authenticated user. Each result includes its target entity and optional subresource, the user-provided proposalNote (also returned as description for compatibility), and the type-specific member of params. Inspect that content before calling accept_or_reject_proposals; never approve proposals based only on their URNs.
list_pending_proposals
Explore a dataset's columns, filtering by keyword. Pass `keywords` unless you need every column — an unfiltered dump of a wide schema is ~5x the cost. Keywords match descriptions and terms too, so they find columns whose names you cannot guess. If you already know the names, get_entities(schema_columns=[...]) answers in one call. Keyword-matched columns return their full description; an unfiltered dump returns each description as a preview, and one ending in a `[+N chars]` marker is truncated — pass that column's name as a keyword to get its full definition. Args: urn: Dataset URN keywords: Keywords to filter schema fields (OR matching). Omit only to page through an entire schema. - Single keyword: Treated as one keyword (NOT split on whitespace). Use for field names or exact phrases. - Multiple keywords: Multiple keywords, matches any (OR logic). - None or empty list: Returns all fields in priority order (same as get_entities). Matches against fieldPath, description (including propagated and AI-inferred documentation), label, tags, and glossary terms. Matching fields are returned first, sorted by match count. limit: Maximum number of fields to return (default: 100) offset: Number of fields to skip for pagination (default: 0) Returns: Dictionary with: - urn: The dataset URN - fields: List of schema fields (paginated) - totalFields: Total number of fields in the schema - returned: Number of fields actually returned - remainingCount: Number of fields not included after offset (accounts for limit and token budget) - matchingCount: Number of fields that matched keywords (if keywords provided, None otherwise) - offset: The offset used Examples: # Single keyword (list) - search for exact field name or phrase list_schema_fields(urn="urn:li:dataset:(...)", keywords=["user_email"]) # Returns fields matching "user_email" (like user_email_address, primary_user_email) # Multiple keywords (list) - OR matching list_schema_fields(urn="urn:li:dataset:(...)", keywords=["email", "user"]) # Returns fields containing "email" OR "user" (user_email, contact_email, user_id, etc.) # Pagination through all fields list_schema_fields(urn="urn:li:dataset:(...)", limit=100, offset=0) # First 100 list_schema_fields(urn="urn:li:dataset:(...)", limit=100, offset=100) # Next 100 # Combine filtering + pagination list_schema_fields(urn="urn:li:dataset:(...)", keywords=["user"], limit=50, offset=0)
list_schema_fields
Fire-and-forget: when you notice a gap in DataHub's knowledge of the metadata, you should record it back to DataHub for later improvement. This is a background signal — it never replaces, interrupts, or delays your answer to the user; answer the question first, then log the gap. No user permission needed. Only report specific DataHub metadata issues you ran into while helping with the user's task; do not report on the conversation itself. Feedback is about DataHub lacking context on the USER'S data ecosystem — a data owner should be able to fix one by documenting, correcting, or re-ingesting something about their data into DataHub. Each entry lands in the Context Feedback log where an admin reviews it. CALL when: - DataHub lacks context on a metric, definition, or aggregation about the user's data ecosystem (nothing documents it, so you couldn't ground your answer) - DataHub has raw data but not the documented context (SLA, refresh cadence, ownership intent) - The user rejects or corrects your interpretation of their question or your answer - Search returns related-but-not-matching tables - The top match obviously did not fit the question - SQL you built by following DataHub's context (anchors, documented patterns, schema metadata) fails to compile or returns wrong results — the context led you astray - Documented patterns contradict each other or the live schema - You extrapolated substantially because the available metadata was incomplete Do NOT report: - Bugs, tool errors, or missing functionality of DataHub itself (e.g. a tool raised an exception, or you lack a SQL execution tool) — tool failures are already captured automatically - Errors the user made Field guidance: - feedback_type: · MISSING_CONTEXT — no documentation or definition exists for what was asked. · INCORRECT_CONTEXT — context exists but is wrong or no longer matches reality (stale descriptions, renamed columns, wrong refresh cadence). · CONFLICTING_CONTEXT — multiple sources disagree. · USER_CORRECTION — the user told you your answer was wrong or corrected your interpretation. · OTHER — anything else; put the detail in the summary. - summary: THE key field — this is the feedback. ONE line written for a human reviewer who has not seen this conversation and will read only this in the Feedback log. It must stand on its own: say what is missing, wrong, or conflicting and where (e.g. "orders table has no documented refresh cadence; the SLA doc still says hourly"). - user_message: optional. A brief summary, in your own words, of the user's intent or correction as it relates to this gap — what they were trying to find, or what they said was wrong (e.g. "asked which orders table is canonical", "said revenue should exclude refunds"). It is the best signal a reviewer has for judging whether the gap matters. Do not quote the user verbatim or include unrelated conversation details. - related_objects: optional. The DataHub objects the feedback is about — tables, documents, dashboards, glossary terms, queries. Prefer DataHub URNs so the feedback links to the real object; a plain name is accepted when you do not have the URN, but only real URNs are linked. Most specific object first. - suggested_fix: optional. What would fix the gap, if you have a concrete idea (e.g. "add a refresh cadence to the orders table description"). Recorded for analysis; a reviewer decides what to do. - summary, user_message, suggested_fix: describe gaps in structural or schema terms only. Never include actual data values, query results, row counts, or sample rows. Returns ``{"recorded": true, "urn": <feedback urn>}``. ``recorded`` means the observation was captured; ``urn`` is null if it could not be stored as an entity. Do not retry on a null urn.
note_metadata_observation
Submit a proposal to create a new glossary term for steward review. Routes creation through the ActionRequest approval workflow — the term is not created until a steward accepts the proposal via accept_or_reject_proposals. Use list_pending_proposals to find the resulting proposal URN.
propose_create_glossary_term
Propose a description update for an asset or column for steward review. The description is not changed until a reviewer accepts the proposal. Use this instead of update_description when the change should go through the proposal inbox. Args: entity_urn: Asset URN to update. description: Proposed description text (markdown supported). column_path: Optional column name for a schema-field description. Omit for asset-level documentation. rationale: Optional note shown to reviewers (proposalNote). Returns: success and the target urn/column. This GraphQL mutation returns a boolean rather than a proposal URN; use list_pending_proposals with proposal_type="UPDATE_DESCRIPTION" to find the ActionRequest.
propose_description
Submit title or content changes to an existing document for review. Supply the proposed fields directly. DataHub creates and manages the internal draft copy; callers never create or pass a draft URN. Omitted title or content fields retain their current values. At least one proposed field is required.
propose_document_edit
Propose assigning a domain to assets for steward review. Domain assignment is entity-level only. The domain is not applied until a reviewer accepts each proposal. Use this instead of set_domains when the change should go through the proposal inbox. Args: domain_urn: Domain URN to propose (e.g. "urn:li:domain:marketing"). entity_urns: Asset URNs to assign to the domain. rationale: Optional note shown to reviewers.
propose_domain
Propose applying glossary terms to assets or columns for steward review. The terms are not applied until a reviewer accepts the resulting ActionRequest. Each asset (or column) gets its own proposal. Use this instead of add_terms when the change should go through the proposal inbox. Args: term_urns: Glossary term URNs to propose (e.g. ["urn:li:glossaryTerm:Customer"]). entity_urns: Asset URNs to apply the terms to. column_paths: Optional column names. Pair one-for-one with entity_urns, or pass one entity URN with multiple columns to target several fields on one dataset. None or an empty string means the whole asset. rationale: Optional note shown to reviewers. Returns: success, proposals (resourceUrn, column_path, proposalUrn), and a message.
propose_glossary_terms
Submit a proposal to change an entity's lifecycle stage. You MUST call list_lifecycle_stages first with the target entity type. Choose a result whose userAssignable field is true. Never guess or hardcode a stage URN. DRAFT is an internal working-copy stage for documents and cannot be proposed for them. stage_urn: a URN returned by list_lifecycle_stages(), or None to propose clearing the current stage. description: optional note explaining the reason for the proposed change. Returns the URN of the created proposal (ActionRequest).
propose_lifecycle_stage
Propose adding owners to assets for steward review. Ownership is entity-level only (not columns). Owners are not applied until a reviewer accepts each proposal. Use this instead of add_owners when the change should go through the proposal inbox. Args: owner_urns: CorpUser or CorpGroup URNs to propose. entity_urns: Asset URNs to assign ownership to. ownership_type: Built-in OwnershipType (TECHNICAL_OWNER, BUSINESS_OWNER, DATA_STEWARD) or a custom ownership type name / URN. rationale: Optional note shown to reviewers.
propose_owners
Propose structured property values on assets or a column for steward review. Values are not applied until a reviewer accepts each proposal. Use this instead of add_structured_properties when the change should go through the proposal inbox. Args: property_values: Map of structured property URN to a list of values. entity_urns: Asset URNs to apply the properties to. column_paths: Optional schema field paths. Pair one-for-one with entity_urns, or pass one entity URN with multiple paths to target several fields on one dataset. None or an empty string means the whole asset. rationale: Optional note shown to reviewers.
propose_structured_properties
Propose applying tags to assets or columns for steward review. The tags are not applied until a reviewer accepts the resulting ActionRequest. Each asset (or column) gets its own proposal. Use this instead of add_tags when the change should go through the proposal inbox. Args: tag_urns: Tag URNs to propose (e.g. ["urn:li:tag:PII"]). entity_urns: Asset URNs to apply the tags to. column_paths: Optional column names. Pair one-for-one with entity_urns, or pass one entity URN with multiple columns to target several fields on one dataset. None or an empty string means the whole asset. rationale: Optional note shown to reviewers. Returns: success, proposals (resourceUrn, column_path, proposalUrn), and a message.
propose_tags
Remove domain assignment from multiple DataHub entities. This tool allows you to unset the domain for multiple entities in a single operation. Useful for removing domain assignments when reorganizing entities or correcting misassignments. Args: entity_urns: List of entity URNs to remove domain from (e.g., dataset URNs, dashboard URNs) Returns: Dictionary with: - success: Boolean indicating if the operation succeeded - message: Success or error message Examples: # Remove domain from multiple datasets remove_domains( entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.old_table,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.deprecated,PROD)" ] ) # Remove domain from dashboards remove_domains( entity_urns=[ "urn:li:dashboard:(urn:li:dataPlatform:looker,old_dashboard,PROD)", "urn:li:dashboard:(urn:li:dataPlatform:looker,temp_dashboard,PROD)" ] ) # Remove domain from mixed entity types remove_domains( entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.temp,PROD)", "urn:li:dataFlow:(urn:li:dataPlatform:airflow,old_pipeline,PROD)", "urn:li:dashboard:(urn:li:dataPlatform:superset,test,PROD)" ] )
remove_domains
Remove one or more owners from multiple DataHub entities. This tool allows you to unassign multiple entities from multiple owners in a single operation. Useful for bulk ownership removal operations like removing owners when they change roles, cleaning up stale ownership, or correcting misassigned ownership. Note: Ownership in DataHub is entity-level only. For field-level metadata, use tags or glossary terms instead. Args: owner_urns: List of owner URNs to remove (must be CorpUser or CorpGroup URNs). Examples: ["urn:li:corpuser:john.doe", "urn:li:corpGroup:data-engineering"] entity_urns: List of entity URNs to remove ownership from (e.g., dataset URNs, dashboard URNs) ownership_type: Optional ownership type to specify which type of ownership to remove. If not provided, will remove ownership regardless of type. Accepts either a built-in OwnershipType enum value (TECHNICAL_OWNER, BUSINESS_OWNER, DATA_STEWARD) or a custom ownership type by its user-facing name (e.g. "Producer"). Custom types are looked up by name (case-insensitive) and must already exist in DataHub. Returns: Dictionary with: - success: Boolean indicating if the operation succeeded - message: Success or error message Examples: # Remove owners from multiple datasets (any ownership type) remove_owners( owner_urns=["urn:li:corpuser:former.employee", "urn:li:corpGroup:old-team"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.customers,PROD)" ] ) # Remove technical owner with specific ownership type remove_owners( owner_urns=["urn:li:corpuser:john.doe"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.customers,PROD)" ], ownership_type=TECHNICAL_OWNER ) # Remove temporary owner from multiple entities remove_owners( owner_urns=["urn:li:corpuser:temp.owner"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.stable_table,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", "urn:li:dashboard:(urn:li:dataPlatform:looker,temp_dashboard,PROD)" ] )
remove_owners
Remove structured properties from multiple DataHub entities. This tool allows you to remove structured property assignments from multiple entities in a single operation. Args: property_urns: List of structured property URNs to remove Example: ["urn:li:structuredProperty:io.acryl.privacy.retentionTime"] entity_urns: List of entity URNs to remove properties from (e.g., dataset URNs, dashboard URNs) Examples: # Remove retention time property from datasets remove_structured_properties( property_urns=["urn:li:structuredProperty:io.acryl.privacy.retentionTime"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.old_data,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.archived,PROD)" ] ) # Remove multiple properties at once remove_structured_properties( property_urns=[ "urn:li:structuredProperty:io.acryl.privacy.retentionTime", "urn:li:structuredProperty:io.acryl.common.businessCriticality" ], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.temp_table,PROD)" ] )
remove_structured_properties
Remove one or more tags from multiple DataHub entities or their column_paths (e.g., schema fields). This tool allows you to untag multiple entities or their columns with multiple tags in a single operation. Useful for bulk tag removal operations like removing deprecated tags, correcting misapplied classifications, or cleaning up governance metadata. Args: tag_urns: List of tag URNs to remove (e.g., ["urn:li:tag:PII", "urn:li:tag:Sensitive"]) entity_urns: List of entity URNs to untag (e.g., dataset URNs, dashboard URNs) column_paths: Optional list of column_path identifiers (e.g., column names for schema fields). Must be same length as entity_urns if provided. Use None or empty string for entity-level tag removal. For column-level tag removal, provide the column name (e.g., "email_address"). Verify that the column_paths are correct and valid via the schemaMetadata. Use get_entity tool to verify. Returns: Dictionary with: - success: Boolean indicating if the operation succeeded - message: Success or error message Examples: # Remove tags from multiple datasets remove_tags( tag_urns=["urn:li:tag:Deprecated", "urn:li:tag:Legacy"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.old_users,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.old_customers,PROD)" ] ) # Remove tags from specific columns remove_tags( tag_urns=["urn:li:tag:PII"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)" ], column_paths=["old_email_field", "deprecated_phone"] ) # Mix entity-level and column-level tag removal remove_tags( tag_urns=["urn:li:tag:Experimental"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.stable_table,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)" ], column_paths=[None, "test_column"] # Remove from whole table and a specific column )
remove_tags
Remove one or more glossary terms (terms) from multiple DataHub entities or their column_paths (e.g., schema fields). This tool allows you to disassociate multiple entities or their columns from multiple glossary terms in a single operation. Useful for bulk term removal operations like correcting misapplied business definitions, updating terminology, or cleaning up metadata. Args: term_urns: List of glossary term URNs to remove (e.g., ["urn:li:glossaryTerm:Deprecated", "urn:li:glossaryTerm:Legacy"]) entity_urns: List of entity URNs to remove terms from (e.g., dataset URNs, dashboard URNs) column_paths: Optional list of column_path identifiers (e.g., column names for schema fields). Must be same length as entity_urns if provided. Use None or empty string for entity-level glossary term removal. For column-level glossary term removal, provide the column name (e.g., "old_field"). Verify that the column_paths are correct and valid via the schemaMetadata. Use get_entity tool to verify. Returns: Dictionary with: - success: Boolean indicating if the operation succeeded - message: Success or error message Examples: # Remove glossary terms from multiple datasets remove_glossary_terms( term_urns=["urn:li:glossaryTerm:Deprecated", "urn:li:glossaryTerm:LegacySystem"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.old_users,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.old_customers,PROD)" ] ) # Remove glossary terms from specific columns remove_glossary_terms( term_urns=["urn:li:glossaryTerm:Confidential"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)" ], column_paths=["old_ssn_field", "legacy_tax_id"] ) # Mix entity-level and column-level glossary term removal remove_glossary_terms( term_urns=["urn:li:glossaryTerm:Experimental"], entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.production_table,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)" ], column_paths=[None, "beta_feature"] # Remove from whole table and a specific column )
remove_terms
Render structured results as a card, table, bar chart, or line chart. Call this after retrieving the factual data. Use a summary card for one key metric, a table for exact multi-column values, and a bar chart for comparing categories or rankings. Use a line chart for values measured over time. Keep the final text response concise and do not repeat every value already shown by the visualization. Args: widget_id: A unique, descriptive identifier for this visualization. spec: The visualization specification to render. Returns: A host-neutral payload that chat and MCP App hosts render as given.
render_visualization
Save or update a STANDALONE document in DataHub's knowledge base. Once saved, a document will be visible to all users of DataHub and to Ask DataHub AI assistant. NOTE: This tool is for creating standalone documents (insights, FAQs, notes, etc.), NOT for updating descriptions on data assets like datasets or dashboards. Use update_description for asset descriptions. WHEN TO USE THIS TOOL: Use this tool when the user explicitly requests to save information: - "Save this for later..." - "Bookmark this.." - "Document this insight.." - "Remember this.." - "Add this to our knowledge base.." - "Create a document about this.." Also SUGGEST using this tool when the user provides valuable information such as: - Useful SQL queries they want to reuse - Decisions about data modeling or architecture - FAQs or common questions about data - Analysis results worth sharing with the team - Corrections or clarifications about data, service, business definitions, etc. ⚠️ IMPORTANT: Before calling this tool, you SHOULD confirm with the user that they want to save this document. Present the title, content summary, and any related assets, and ask for their approval before proceeding. Do not attempt to save information that would be private or user-specific. This tool persists insights, decisions, FAQs, and other contextual information as documents in DataHub. Documents are organized hierarchically: - Under a configurable parent folder (default: "Shared" for global context) - Optionally grouped by the user who authored them UPSERT BEHAVIOR: - If `urn` is NOT provided: Creates a NEW document with a unique URN - If `urn` IS provided: Updates the EXISTING document with that URN IMPORTANT USAGE GUIDELINES: - Always confirm with the user before saving - Provide a clear summary of what will be saved - Ask if the user wants to proceed with creating/updating the document REQUIRED PARAMETERS: document_type: The type of document being saved. For example: - "Insight": Data insights or discoveries - "Decision": Documented decisions with rationale - "FAQ": Frequently asked questions and answers - "Analysis": Data analysis findings - "Summary": Summaries of complex information - "Recommendation": Suggested actions or improvements - "Note": General notes or observations title: A descriptive title for the document. - Example: "Sales Data Quality Issues - Q4 2024" - Example: "Decision: Deprecating Legacy Customer Table" content: The full content of the document (supports markdown formatting). - Can include headers, lists, code blocks, tables, etc. - Example: "## Summary\n\nThe orders table shows 15% null values..." OPTIONAL PARAMETERS: urn: The URN of an existing document to update. - ONLY use after a search_documents or get_entity call returns a document URN - Example: "urn:li:document:agent-insight-abc123" - If not provided, a new document is created with a unique URN - If provided, the existing document is updated (upsert operation) topics: List of topic tags for categorization and discovery (like a word cloud). - These become searchable tags in DataHub that users can click to find related documents - Example: ["data-quality", "customer-data", "Q4-2024"] - Example: ["high-priority", "sales", "email", "null-values"] related_documents: URNs of related documents. - Example: ["urn:li:document:agent-insight-sales-abc123"] - Creates links between related knowledge related_assets: URNs of related data assets (tables, dashboards, etc). - Example: ["urn:li:dataset:(urn:li:dataPlatform:snowflake,db.orders,PROD)"] - Links the document to specific data assets in the catalog - Users can then see this document when viewing those assets Returns: Dictionary with: - success: Boolean indicating if the operation succeeded - urn: The URN of the created/updated document - message: Success or error message - author: The user who authored the document (if available) RECOMMENDED WORKFLOW: 1. Gather information you want to save publicly 2. Present a summary to the user: "I'd like to save the following insight to DataHub: - Title: High Null Rate in Customer Emails - Type: Insight - Related to: customers table Would you like me to save this?" 3. Only call save_document after user confirms EXAMPLE USAGE: 1. Create a new insight (after user confirmation): save_document( document_type="Insight", title="High Null Rate in Customer Emails", content="## Finding\n\n23% of customer records have null email...", topics=["data-quality", "customer-data", "email", "high-severity"], related_assets=["urn:li:dataset:(urn:li:dataPlatform:snowflake,customers,PROD)"] ) 2. Update an existing document (after finding it via search_documents): save_document( urn="urn:li:document:agent-insight-abc123", # From search_documents result document_type="Insight", title="High Null Rate in Customer Emails (Updated)", content="## Finding\n\nUpdated: Now 18% of customer records have null email...", topics=["data-quality", "customer-data", "email", "resolved"] ) 3. Document a decision: save_document( document_type="Decision", title="Migrating to New Production Database", content="## Decision\n\nWe will migrate to v2 schema...\n\n## Rationale\n...", topics=["architecture", "data-model", "migration", "approved"] )
save_document
Search across DataHub entities using structured full-text search. Results are ordered by relevance and importance - examine top results first. SEARCH SYNTAX: - Structured full-text search - **always start queries with /q** - **Recommended: Use + operator for AND** (handles punctuation better than quotes) - Supports full boolean logic: AND (default), OR, NOT, parentheses, field searches - Examples: * /q user+transaction -> requires both terms (better for field names with _ or punctuation) * /q point+sale+app -> requires all terms (works with point_of_sale_app_usage) * /q wizard OR pet -> entities containing either term * /q revenue* -> wildcard matching (revenue_2023, revenue_2024, revenue_monthly, etc.) * /q tag:PII -> search by tag name * /q "exact table name" -> exact phrase matching (use sparingly) * /q (sales OR revenue) AND quarterly -> complex boolean combinations - Fast and precise for exact matching, technical terms, and complex queries - Best for: entity names, identifiers, column names, or any search needing boolean logic PAGINATION: - num_results: Number of results to return per page (max: 50) - offset: Starting position in results (default: 0) - Examples: * First page: offset=0, num_results=10 * Second page: offset=10, num_results=10 * Third page: offset=20, num_results=10 FACET EXPLORATION - Discover metadata without returning results: - Set num_results=0 to get ONLY facets (no search results) - Facets show ALL tags, glossaryTerms, platforms, domains used in the catalog - Example: search(query="*", filter="entity_type = dataset", num_results=0) -> Returns facets showing all tags/glossaryTerms applied to datasets - Use this to discover what metadata exists before doing filtered searches TYPICAL WORKFLOW: 1. Facet exploration: search(query="*", filter="entity_type = dataset", num_results=0) -> Examine tags/glossaryTerms facets to see what metadata exists 2. Filtered search: search(query="*", filter="tag = urn:li:tag:pii", num_results=30) -> Get entities with specific tag using URN from step 1 3. Get details: Use get_entities() on specific results FILTER SYNTAX: SQL-like WHERE clause over the fields below. Supports AND, OR, NOT, parentheses, IN (a, b), comparison (>, >=, <, <=), and IS [NOT] NULL. e.g. entity_type = dataset AND env = PROD AND (platform = snowflake OR platform = bigquery) SUPPORTED FILTER FIELDS: - entity_type: dataset, dashboard, chart, corp_user, corp_group, dataProduct, etc. - entity_subtype (or subtype): Table, View, Model, etc. - platform: snowflake, bigquery, looker, tableau, etc. - domain, container, tag, glossary_term, owner: full URN (urn:li:...), found by searching entity_type = domain / container / tag / glossaryTerm / corp_user - env: PROD, DEV, STAGING (only use if explicitly requested) - status: NOT_SOFT_DELETED, ALL, ONLY_SOFT_DELETED - deprecated, hasActiveIncidents, hasFailingAssertions: true or false - columnCount: number of columns (from dataset profiling) - lastModifiedAt, createdAt: when the asset was last changed / created in the source system (datasets, dashboards, charts, pipelines, containers, etc.) - lastOperationTime: time of the last write seen in the source's audit logs (datasets only) Compare time fields to an ISO date, e.g. lastModifiedAt >= 2026-01-31 STRUCTURED PROPERTY FILTERS: use structuredProperties.<qualifiedName> as the field name, where qualifiedName is the ID portion of the structured property URN, e.g. structuredProperties.io.acryl.privacy.retentionTime = 30. Values are CASE-SENSITIVE; some properties have fixed allowed values (e.g. "Published"), so use get_entities on the property URN to check definition.allowedValues. Double-quote any value containing spaces, =, or parentheses: subtype = "Materialized View" subtype IN ("Materialized View", Table) SEARCH STRATEGY EXAMPLES: - /q customer+behavior -> finds tables with both terms (works with customer_behavior fields) - /q customer OR user -> finds tables with either term - /q (financial OR revenue) AND metrics -> complex boolean logic SORTING - Order results by specific fields: - sort_by: Field name to sort by (optional) - sort_order: "desc" (default) or "asc" Available sort fields: - queryCountLast30DaysFeature: Number of queries in last 30 days - rowCountFeature: Table row count - sizeInBytesFeature: Table size in bytes - writeCountLast30DaysFeature: Number of writes/updates in last 30 days - lastOperationTime: Time of the last write seen in the source's audit logs - lastModifiedAt: When the asset was last changed in the source system - createdAt: When the asset was created in the source system Sorting examples: - Most queried datasets: search(query="*", filter="entity_type = dataset", sort_by="queryCountLast30DaysFeature", num_results=10) - Largest tables: search(query="*", filter="entity_type = dataset", sort_by="sizeInBytesFeature", num_results=10) - Smallest tables first: search(query="*", filter="entity_type = dataset", sort_by="sizeInBytesFeature", sort_order="asc", num_results=10) - Most recently updated: search(query="*", filter="entity_type = dataset", sort_by="lastOperationTime", sort_order="desc", num_results=10) - Newest dashboards: search(query="*", filter="entity_type = dashboard", sort_by="createdAt", sort_order="desc", num_results=10) Note: If sort_by is not provided, search results use default ranking by relevance and importance. When using sort_by, results are strictly ordered by that field.
search
Search for documents stored in the customer's DataHub deployment. These are the organization's own documents (runbooks, FAQs, knowledge articles) ingested from sources like Notion, Confluence, etc. - not DataHub documentation. Returns document metadata WITHOUT content to keep responses concise. To read a document's full body, pass its URN to grep_documents(). HYBRID SEARCH (recommended for natural language queries): When both query and semantic_query are provided, runs keyword and semantic searches in parallel, merges results intelligently, and applies AI relevance reordering so the most relevant result is first: - Results are deduplicated by URN - Each result includes searchType: "keyword", "semantic", or "both" - Results appearing in both searches are high-confidence matches - Each result carries a ``rerankScore`` (0-1) — higher means more relevant to the query; useful as a confidence signal when deciding how many results to actually look at Example: search_documents( query="/q kubernetes+deployment", semantic_query="how do I deploy applications to kubernetes cluster" ) KEYWORD SEARCH (query parameter): - Always prefix with /q — a bare query matches as one exact phrase - Inside /q: + for AND (/q deployment+guide), also OR, NOT, (), "quoted phrase" - `*` = browse-all for facets / filter-only; keyword-only, not with semantic_query SEMANTIC SEARCH (semantic_query parameter): - Uses AI embeddings to find conceptually related documents - Best for: natural language questions, finding related topics - Example: "how to deploy" finds deployment guides, CI/CD docs, release runbooks FILTER SYNTAX: SQL-like WHERE clause over the fields below. Supports AND, OR, NOT, parentheses, and IN (a, b). e.g. platform IN (notion, confluence) AND tag = urn:li:tag:critical SUPPORTED FILTER FIELDS: - subtype (or entity_subtype): values vary by deployment and are CASE-SENSITIVE; enumerate them with the facet call below, don't guess - platform: notion, confluence, datahub, etc. - domain, tag, glossary_term, owner: full URN required IMPORTANT: domain, tag, glossary_term, and owner require full URN format (urn:li:...). Use the facet discovery call below to find valid URNs, then use the exact URN from the results. Double-quote any value containing spaces, =, or parentheses: tag = "urn:li:tag:my tag" subtype IN ("Two Words", OneWord) PAGINATION: - num_results: Number of results per page (max: 50) - offset: Starting position (default: 0) FACET DISCOVERY: - Set num_results=0 to get ONLY facets (no results) - Pass query="*" (optionally with a filter) to enumerate the values that actually exist: the "Sub Type" facet (field typeNames) lists valid subtypes, alongside platform, domain and tags - Do this before filtering on subtype. A guessed subtype is a valid filter that matches nothing, so it returns zero results rather than an error EXAMPLE WORKFLOWS: 1. Hybrid search for deployment docs: search_documents( query="/q kubernetes+deployment", semantic_query="how to deploy applications to production" ) 2. Keyword-only search filtered by platform: search_documents(query="deployment", filter="platform = notion") 3. Discover valid subtypes / platforms before filtering: search_documents(query="*", num_results=0) → typeNames facet lists the subtypes that exist, with counts 4. Find engineering team's critical docs: search_documents(query="*", filter="domain = urn:li:domain:engineering AND tag = urn:li:tag:critical") 5. Find all docs from Notion or Confluence: search_documents(query="*", filter="platform IN (notion, confluence)")
search_documents
Set domain for multiple DataHub entities. This tool allows you to assign a domain to multiple entities in a single operation. Useful for organizing datasets, dashboards, and other entities into logical business domains. Note: Domain assignment in DataHub is entity-level only. Each entity can belong to exactly one domain. Setting a new domain will replace any existing domain assignment. Args: domain_urn: Domain URN to assign (e.g., "urn:li:domain:marketing") entity_urns: List of entity URNs to assign to the domain (e.g., dataset URNs, dashboard URNs) Returns: Dictionary with: - success: Boolean indicating if the operation succeeded - message: Success or error message Examples: # Set domain for multiple datasets set_domains( domain_urn="urn:li:domain:marketing", entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.campaigns,PROD)", "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.customers,PROD)" ] ) # Set domain for dashboards set_domains( domain_urn="urn:li:domain:finance", entity_urns=[ "urn:li:dashboard:(urn:li:dataPlatform:looker,revenue_dashboard,PROD)", "urn:li:dashboard:(urn:li:dataPlatform:looker,expense_dashboard,PROD)" ] ) # Set domain for mixed entity types set_domains( domain_urn="urn:li:domain:engineering", entity_urns=[ "urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.logs,PROD)", "urn:li:dataFlow:(urn:li:dataPlatform:airflow,etl_pipeline,PROD)", "urn:li:dashboard:(urn:li:dataPlatform:superset,metrics,PROD)" ] )
set_domains
Set or clear the lifecycle stage on an entity (glossary term, document, etc.). IMPORTANT: You MUST call list_lifecycle_stages first with the target entity type. Choose a result whose userAssignable field is true. Never guess or hardcode lifecycle stage URNs. stage_urn: a URN returned by list_lifecycle_stages() (e.g. "urn:li:lifecycleStageType:DRAFT"), or None to clear the stage and publish/activate the entity. Passing None makes the entity visible in default search results. Invalid transitions are silently reverted by the server — check allowedPreviousStages from list_lifecycle_stages() to verify the transition is permitted.
set_lifecycle_stage
Update description for a DataHub entity or its column (e.g., schema field). This tool allows you to set, append to, or remove a description for an entity or its column. Useful for documenting datasets, containers, charts, dashboards, data flows, data jobs, ML models, ML model groups, ML feature tables, ML primary keys, tags, glossary terms, glossary nodes, domains, and schema fields. Args: entity_urn: Entity URN to update description for (e.g., dataset URN, container URN) operation: The operation to perform: - "replace": Replace the existing description with the new one (default) - "append": Append the new description to the existing one - "remove": Remove the description (description parameter not needed) description: The description text to set or append (supports markdown formatting). Required for "replace" and "append" operations, ignored for "remove". column_path: Column_path identifier (e.g., column name for schema field). Optional for all entity types (use None for entity-level descriptions). For column-level descriptions, provide the column name (e.g., "customer_email"). Verify that the column_path is correct and valid via the schemaMetadata. Use get_entity tool to verify. Returns: Dictionary with: - success: Boolean indicating if the operation succeeded - urn: The entity URN - column_path: The column path (if applicable) - message: Success or error message Examples: # Update description for a container (entity-level) update_description( entity_urn="urn:li:container:12345", operation="replace", description="Production data warehouse" ) # Update description for a dataset update_description( entity_urn="urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", operation="replace", description="User's table", ) # Update description for a dataset field (column-level) update_description( entity_urn="urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", operation="replace", description="User's primary email address", column_path="email" ) # Append to existing field description update_description( entity_urn="urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", operation="append", description=" (PII)", column_path="email" ) # Remove field description update_description( entity_urn="urn:li:dataset:(urn:li:dataPlatform:snowflake,db.schema.users,PROD)", operation="remove", column_path="old_field" )
update_description
Partial update of an existing Document — send the diff, not the whole body. Use this for any edit to an existing document; ``save_document`` is for creation and full-replace only. Commands: - ``update_content`` — apply one or more ``{old_str, new_str, replace_all_matches?}`` edits to the body. Fetch the doc first (``get_entities``) so ``old_str`` matches exactly. Multiple matches fail unless ``replace_all_matches=true``. - ``replace_content`` — overwrite the body with ``new_content``. Requires ``allow_deleting_content=true`` if the doc currently has content. - ``append_content`` — concatenate ``new_content`` to the end of the body. - ``update_properties`` — change ``title``, ``topics``, ``related_documents``, ``related_assets``, or ``status`` (``"PUBLISHED"`` / ``"UNPUBLISHED"``) without touching the body. Honors the same ``RESTRICT_SAVE_DOCUMENT_UPDATES_TO_SHARED_FOLDER`` guardrail as ``save_document``. Confirm with the user before ``replace_content`` or publishing. Examples: update_document( urn="urn:li:document:my-doc", command="update_content", content_edits=[{ "old_str": "## ARR\nARR is defined as ...", "new_str": "## ARR\nARR = MRR × 12. Excludes one-time fees ...", }], ) update_document( urn="urn:li:document:my-doc", command="update_properties", properties={"status": "PUBLISHED"}, )
update_document
DataHub Cloud ChatGPT Plugin FAQ
How the directory, categories and Discoverability Score work.
Read the methodologyHow do I improve DataHub Cloud's ChatGPT Plugin discoverability?
The levers are the listing surface agents actually read: names, descriptions, keywords, tool metadata, and registry health. Which lever matters depends on where discovery breaks, which is what continuous measurement shows.
What are DataHub Cloud alternatives on ChatGPT?
As of 2026-10-10, DataHub Cloud competes with Algolia Productivity, Analytics Brain, Archiveye, Atlan, BitsWeave, BuildBetter.ai, Business Helper, CERTENTIC Bridge and 45 more in ChatGPT Enterprise Knowledge Search & AI Context Layer, ranked by public Discoverability Score.
Where is this profile measured?
This profile uses the geography attached to the latest public registry snapshot: US. Locale tags are intentionally omitted.