DOT Databricks Overview
Capabilities
The Department of Transportation (DOT) Enterprise Databricks platform provides a unified, scalable data intelligence environment built on AWS. Designed to streamline data engineering, collaborative data science, and governance, the platform enables teams to securely build and deploy data pipelines and analytics models. Operating across a structured multi-environment architecture, DOT Databricks ensures rigorous separation of duties while offering high-performance processing capabilities for the department’s most demanding data challenges.
Current Projects
OST
CISO
CFAR
S98 Pilot
OST-B
Data Team
CFO Data Team
Grants
FHWA
Economic Investment Strategies
Highway Safety Information System
National Highway Construction Cost Index Reports
Pathway to Advancing Novel Data Analytics
Freight AI
FRA
Geospatial Analytics
Safety ETL
FTA
Enterprise Data Lakehouse
Office of the Chief Data Officer
NHTSA
Crash Data Acquisition Network, Infrastructure Investment and Jobs Act
Enterprise Data Management Analytics Services
PHMSA
DataMart Modernization
Volpe
Data Warehouse
EPA HWTA
Highway Economic Requirements System (HERS) Cost Matrix
FMCSA PORT
NPS Traveler Info
Surface Transportation Bureau
US Coast Guard, Vessel General Permit
💬Questions?
Email Shyla Morisetty
Desk: 202.924.4019
🌐 Environment Architecture
To ensure enterprise-grade stability and strict lifecycle management, the DOT Databricks (AWS) accounts are structured into four distinct, isolated workspaces:
Dev (Development): The primary sandbox for engineers and data scientists to build, experiment, and write code.
Test: A dedicated environment for quality assurance, code validation, and integration testing.
Stage (Staging): A pre-production replica used to validate data pipelines and performance against production-like data scales.
Prod (Production): The highly controlled, operational environment hosting finalized, automated pipelines and business-critical data assets.
🎯 Our Vision
Our vision is to establish Databricks as the premier lakehouse foundation for the DOT, breaking down data silos across modes. By combining data engineering, data warehousing, and machine learning into a single collaborative workspace, we aim to accelerate time-to-insight for transportation safety, infrastructure, and policy analytics.
🔒 Security & Data Governance
Data within DOT Databricks is rigorously secured and governed through a structured Unity Catalog hierarchy tailored around data "Projects":
Data Isolation: Each Databricks Project is provisioned with a dedicated Main Catalog to cleanly isolate its schemas and tables, with the flexibility to add additional catalogs as the project scales.
Role-Based Access Control (RBAC): Access is strictly managed via specialized user groups, ensuring appropriate permissions per project:
Admins Group: Responsible for catalog management, schema creation, and configuration.
Engineers Group: Empowered to build pipelines, manage compute clusters, and transform data.
Custom Groups: Provisioned on-demand (e.g., Data Analysts, Read-Only Consumers) to meet specific project collaboration needs.
📊 Data & External Integration
DOT Databricks acts as a centralized data hub, allowing teams to seamlessly connect to distributed enterprise data assets without complex migration friction:
External Volumes: Securely mount cloud object storage by establishing direct connections to dedicated Amazon S3 buckets for raw file ingestion, staging, and unstructured data analysis.
Foreign Catalogs: Connect directly to external operational databases, data warehouses, and legacy storage environments using federated queries to analyze data right where it lives.
🛠️ Enterprise Support & Onboarding
The DOT Databricks Platform Team supports your project throughout its development lifecycle—from provisioning and catalog creation to performance tuning and production deployment.