Skip to main content

Assign and run a crawler

Where: sidebar → Crawler. The page heading is Data Source Crawler.

A crawler reads structural metadata from a source. It is not a data-ingestion pipeline.

Before you start

You need a verified connection supported by metadata discovery and permission to manage workspace resources. The source account must be allowed to read the relevant metadata.

Assign a source

Select Assign database. Choose the Verified data connection and give the crawler a meaningful name, such as training-orders-metadata.

Choose the schedule

Set the Daily time and Timezone. Verify the timezone with the source owner, then select Assign crawler. Scheduling a metadata scan does not schedule pipeline ingestion.

Run the first scan

Select Run now. A queued response confirms submission, not completion. Open View runs and wait for the run to finish.

Inspect what was discovered

Open Source Schema and select Browse schema for this crawler. Check the snapshot timestamp and confirm that expected relations and columns appear.

Keep the scan current

  • Use Edit to change the assignment settings offered by the form.
  • Use Pause and Resume to control the crawler schedule.
  • Use Run now after a known source schema change when appropriate.
  • Remove removes the crawler assignment; its confirmation states that run history is preserved.

If the newest scan fails, an older successful snapshot may still be visible. Read the warning and timestamp before treating that metadata as current.

If no schemas appear

Check the saved connection, source metadata permissions, run outcome, and connector support. If the run is still queued, ask the platform operator to inspect the crawler service. Do not repeatedly create new assignments to work around a failed run.

Continue with Browse Source Schema.