Agent Skills · Cursor · Databricks Genie
Enterprise Data-Engineering Skills for AI Agents
Overview
This project packages the hardest, most repetitive parts of data engineering into a library of enterprise-level agent skills — structured task playbooks that an AI agent can execute deterministically instead of improvising. Each skill encodes the standards, guardrails, and step-by-step procedure for one class of work, so an agent produces consistent, reviewable output whether it is ingesting a new feed, validating a schema, or parsing an EDI file.
The skills are surfaced to agents across multiple runtimes. In Cursor they guide code-writing and review workflows; in Databricks Genie they back natural-language operations over governed data, letting analysts trigger ingestion, validation, and transformation tasks in plain language. Because the same skill definitions drive every runtime, teams get one source of truth for how data-engineering automation should behave — no per-tool drift, no re-implementing the same prompt logic on each surface.
Business value comes from standardization and safety. The library covers data ingestion, data validation, data quality checks, parsing of complex files, interoperability and data-sufficiency checks, EDI parsing, and data transformations. Skills are versioned and centrally governed, so a fix or policy change is authored once and rolled out everywhere. Quality gates and data-sufficiency checks are built into the playbooks themselves, so automation fails loudly on bad input rather than silently propagating it downstream.
Use Cases
Data ingestion
A playbook for onboarding new feeds — source discovery, landing-zone conventions, idempotent loads, and audit columns — so ingestion is repeatable, traceable, and safe to re-run.
Data validation
Schema, type, and constraint validation with explicit pass/fail gates that block malformed records from promoting past the landing layer.
Data quality checks
Standardized completeness, uniqueness, freshness, and referential-integrity checks with measured baselines and clear, actionable failure reporting.
Parsing complex files
Structured extraction from messy real-world files — nested spreadsheets, multi-section documents, irregular layouts — into typed, canonical records.
Interoperability & data-sufficiency checks
Interoperability rule checks plus data-sufficiency gating: confirm required fields and coverage exist before a downstream process is allowed to run.
EDI parsing agents
EDI / X12 parsing that decomposes transaction sets into structured, validated records for healthcare and B2B data exchange.
Data transformations
Canonical transformation patterns — mapping, normalization, and derivations — expressed as reviewable, reusable steps rather than one-off scripts.
Tech Stack
- Agent skills (structured task playbooks) consumed by Cursor and Databricks Genie
- Databricks Genie — natural-language data operations backed by governed skills
- Data ingestion, validation, and data-quality-check skill packs
- Complex-file parsing and EDI / X12 parsing skills
- Interoperability and data-sufficiency gating
- Canonical data-transformation patterns
- Central versioning & governance for one source of truth
Technical Implementation
Author skills as governed playbooks
Each skill is a versioned definition encoding the standard, guardrails, and step-by-step procedure for one class of data-engineering work. Skills are reviewed once and stored centrally so behavior is consistent and auditable.
Publish to agent runtimes
The same skill catalog is exposed to Cursor for code-writing and review workflows and to Databricks Genie for natural-language data operations, so every runtime executes identical, approved procedures.
Enforce quality & sufficiency gates
Validation, data-quality, and data-sufficiency checks run inside the skills themselves. Automation halts on bad or incomplete input and reports the failing gate rather than propagating errors downstream.
Version, roll out, and track
Fixes and policy changes are authored once and rolled out everywhere through versioned skill updates, giving teams change tracking and a single governed source of truth for agent-driven automation.