# Virtual Data Assistant

## Overview

The Virtual Data Assistant (VDA) is an innovative and intelligent software tool designed to revolutionize the way data analysts perform their data analytics job. It leverages the power of Assistive Analytics and Generative AI to streamline workflows, enhance productivity, and empower analysts to make data-driven decisions with ease.

<br>

{% embed url="<https://www.youtube.com/watch?v=GrREiSlEkx8>" %}

## Quick links

{% content-ref url="/pages/C3FGrz431UglZ5YPG6DJ" %}
[What we do](/overview/what-we-do)
{% endcontent-ref %}

{% content-ref url="/pages/FXvKOLRLRNTSvZDm9PBR" %}
[Features](/overview/features)
{% endcontent-ref %}


# What we do

&#x20;Virtual Data Assistant is a game-changer for data analysts, empowering them to perform complex data analytics tasks with ease and efficiency. By harnessing the capabilities of Assistive Analytics and Generative AI, the VDA revolutionizes the data analysis workflow, fostering a data-driven culture and propelling organizations towards data-driven success.

**VDA Enhances Data Analysis Workflow:**

1. **Simplified Data Modeling:**
   * Assistive templates accelerate data modeling, enabling analysts to focus on refining insights rather than getting lost in code development.
   * Automatically generated metadata documentation ensures transparency and enhances collaboration among team members.
2. **Efficient Data Cataloging:**
   * Data connectors and metadata extraction streamline the process of data cataloging, empowering analysts to discover and leverage relevant data assets quickly.
   * A centralized view of metadata fosters data governance and aids in maintaining data quality.
3. **Intuitive Dashboard Creation:**
   * NLP-powered dashboard and chart building democratize data visualization, making it accessible to a broader audience within the organization.
   * Visual representations of data drive better decision-making and communication of key insights.
4. **Seamless Workbooks for Analysis:**
   * Document, Query, and Expectation Workbooks collectively expedite data analysis tasks, ensuring analysts can efficiently perform their duties with a reduced risk of errors.
   * VDA promotes iterative and data-driven decision-making, facilitating a culture of continuous improvement.

##


# Features

Below are some of the features we are continuously adding new features so do keep checking

## **Key Features and Benefits:**

1. **Data Cataloging with Metadata Extraction:**
   * Connectors to diverse data sources allow the VDA to perform efficient data cataloging. It automatically extracts metadata from these sources, providing analysts with a comprehensive view of available data assets.
   * The ability to understand data context enables analysts to make informed decisions and facilitates effective collaboration across the organization.
2. **Assistive Templates for Data Modeling:**
   * The VDA offers a collection of intelligent data modeling templates that can generate code for various analytical tasks. These templates significantly reduce development time and minimize the need for manual coding.
   * Additionally, analysts can leverage the VDA to generate metadata documentation automatically, ensuring that the analytical process remains transparent and well-documented.
3. **NLP-Powered Dashboard and Chart Creation:**
   * With the VDA's natural language processing (NLP) capabilities, analysts can effortlessly create interactive dashboards and charts. By simply providing key performance indicators (KPIs) in plain language, the VDA transforms them into visually appealing and insightful data representations.
   * This feature simplifies data visualization, making it accessible to analysts with varying levels of technical expertise.
4. **Workbook Collection:**
   * Document Workbook: Allows analysts to query documents using natural language. The VDA interprets user queries, swiftly retrieving relevant information from documents and presenting it in a structured manner.
   * Query Workbook: Facilitates SQL query writing for analysts through text-to-SQL functionality. Analysts can express their intentions in natural language, and the VDA automatically converts it into SQL code, ensuring accurate and efficient querying.
   * Expectation Workbook: Empowers analysts to perform data validation effortlessly. The VDA provides a range of expectations, enabling users to validate data against predefined criteria, enhancing data quality and accuracy.

<br>

##

## &#x20;

##


# Data Sources

A Datasource is an entity within the Virtual Data Assistant that serves as a container for a collection of metadata. Metadata refers to information about datasets, such as data source location, schema, data types, and other relevant properties. Essentially, a Datasource is like a virtual folder that groups related datasets together, making it easier for users to manage and access data efficiently.

**Datasource creation using Connectors**

<figure><img src="https://4133454587-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FfOPtN4DPW3XVXKtINi7z%2Fuploads%2F1Guc77yQEtdaPGOg6VYR%2FScreen%20Shot%202023-07-29%20at%205.55.49%20PM.png?alt=media&#x26;token=6e4e9b96-511d-4876-a8ab-eaa1de3d2a11" alt=""><figcaption><p>Create a datasource</p></figcaption></figure>

Connectors are modules or plugins that establish connections to specific data sources or databases. VDA offers multiple connectors to popular datasources e.g PostgreSQL, MySQL etc. When users select a connector and provide the necessary connection details, it establishes a link to the data source, allowing access to the data within that source.

**Searching Datasets within a Datasource**

<figure><img src="https://4133454587-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FfOPtN4DPW3XVXKtINi7z%2Fuploads%2FB4qLhMVdMeR2phRCMg5B%2FScreen%20Shot%202023-07-29%20at%2010.42.39%20PM.png?alt=media&#x26;token=0f0657a2-7078-48a9-9442-ef4f9c5b17ed" alt=""><figcaption></figcaption></figure>

Once datasets are created and organized within a Datasource, users can perform searches to find specific entities or datasets within that Datasource. This search functionality simplifies data discovery, especially in cases where multiple datasets are stored within the same Datasource. An user can search using specific filters to find appropriate datasets.

**List of Available Connectors (Grouped by Category)**

**Database Connectors (54)**\
Amazon Athena, Azure SQL, BigQuery, BigTable, BurstIQ, Cassandra, Clickhouse, Cockroach, Couchbase, Custom Database, Databricks, Datalake, Db2, DeltaLake, DomoDatabase, Doris, Dremio, Druid, DynamoDB, Epic, Exasol, Glue, Greenplum, Hive, Iceberg, Impala, Informix, MariaDB, Microsoft Access, Microsoft Fabric, MongoDB, MSSQL, MySQL, Oracle, PinotDB, Postgres, Presto, Redshift, SAS, SQLite, SSAS, Salesforce, SAP ERP, SAP HANA, ServiceNow, SingleStore, Snowflake, StarRocks, Synapse, Teradata, Timescale, Trino, Unity Catalog, Vertica

**Dashboard Connectors (20)**\
Custom Dashboard, Domo Dashboard, Grafana, Hex, Lightdash, Looker, Metabase, MicroStrategy, Mode, PowerBI, PowerBI Report Server, QlikCloud, QlikSense, QuickSight, Redash, Sigma, Ssrs, Superset, Tableau, ThoughtSpot

**Pipeline Connectors (25)**\
Airbyte, Airflow, Backend, Custom Pipeline, DBT Cloud, Dagster, Data Factory, Databricks Pipeline, Domo Pipeline, Fivetran, Flink, Glue Pipeline, Kafka Connect, Kinesis Firehose, Matillion, Microsoft Fabric Pipeline, Mulesoft, Nifi, OpenLineage, SSIS, Snowplow, Spark, Spline, Stitch, Wherescape

**Messaging Connectors (5)**\
Custom Messaging, Kafka, Kinesis, Pulsar, Redpanda

**LLM (Large Language Model) Connectors (8)**\
Anthropic, Azure OpenAI, Bedrock, Custom LLM, HuggingFace, Ollama, OpenAI, VertexAI

**Storage Connectors (4)**\
Adls, Custom Storage, GCS, S3

**Search Connectors (3)**\
Custom Search, Elasticsearch, OpenSearch

**Drive Connectors (4)**\
Custom Drive, Google Drive, SFTP, SharePoint

**ML Model Connectors (5)**\
Custom ML Model, Mlflow, SageMaker, Sklearn, VertexAI

**Metadata Connectors (7)**\
Alation, AlationSink, Amundsen, Atlas, Collibra, MetadataES, OpenMetadata

**API & Security Connectors (2)**\
REST API, Ranger (Security)


# Datasets

Datasets are the lifeblood of data analysis, and   Virtual Data Assistant (VDA) empowers users with a feature-rich environment to harness the full potential of data. Datasets within the VDA serve as powerful entities that showcase data as a product, providing valuable insights and enabling data-driven decision-making. Let's delve into the key features of Datasets:

Types of Datasets

* Live

&#x20;Live datasets refer to datasets that are directly connected to real  data sources. These datasets are  updated and reflect the most current information available from the connected data sources. VDA establishes and maintains live connections to these data sources, ensuring that any changes or updates in  metadata of source data are immediately reflected in the corresponding dataset.

* Custom

Custom datasets, on the other hand, are datasets that are created and defined by users to meet specific analytical requirements. Unlike live datasets, custom datasets are not directly linked to real-time data sources. Instead, users define the data they want to include in the dataset and may apply data transformations to manipulate and refine the data to suit their analytical needs.

<br>

**1. Metadata, Overview, and Tags:**&#x20;

{% embed url="<https://www.loom.com/share/134cf121f45f4ee29d6524488a6ea07b>" %}

Datasets are accompanied by comprehensive metadata, offering essential information about the data they contain. This metadata includes a list of columns, descriptions for each column, and data types, providing users with a clear understanding of the dataset's structure and content. Additionally, an Overview section offers a concise summary of the dataset's key characteristics and significance.

To further enhance organization and accessibility, users can add tags to datasets. Tags serve as labels or keywords, making it easier to search and categorize datasets effectively.

**2. Preview - See Live Data**

{% embed url="<https://www.loom.com/share/6932d5f8efcd48eeb665fe70fe2fe446?sid=5a3f68ce-7063-4fe1-8236-4cd1c5d510ae>" %}

Datasets come to life with the "Preview" feature, offering users a glimpse of the actual data contained within. This live data preview allows analysts to quickly assess the dataset's contents before diving into in-depth analyses. It also aids in identifying any potential data issues or anomalies that may require attention.

**3. Stories - Fostering Collaborative Discussions:**&#x20;

{% embed url="<https://www.loom.com/share/466bff30b64841a597b4c103dd5b66da?sid=6df492dd-23cd-40f1-8e8e-b8b719764fdb>" %}

The "Stories" section provides a collaborative platform for users to engage in discussions about datasets. It serves as a dynamic forum where data analysts, stakeholders, and team members can share insights, ask questions, and exchange valuable perspectives related to the dataset. Additionally, users can upload important documents, transforming the Stories section into a central hub for discussions and knowledge-sharing.

**4. Transformations - Empowering Custom Datasets**

{% embed url="<https://www.loom.com/share/311f7a2b249a4d17b09bf2474521f408?sid=a1ab490a-9cac-45ba-a767-e8c653753e5b>" %}
Create Transformations
{% endembed %}

For custom datasets, the VDA offers the "Transformations" feature. This powerful tool enables users to define data transformations using templates, facilitating data modeling and customization. Analysts can upload transformation logic and document the process, streamlining the preparation and refinement of custom datasets.

**5. Publish and Generate SQL - Effortless Code Generation:**&#x20;

{% embed url="<https://www.loom.com/share/2086acbcc07d43f0b6029b44c5cdb7a7?sid=a73eef3f-0983-44c1-955c-8db73bffd648>" %}

Users can publish versions for transformations and also utilize the generate SQL option to generate code for the defined transformations .

The "Generate SQL" feature leverages the power of OpenAI to automate code generation based on the logic defined in transformations. With just a few clicks, users can effortlessly convert their transformation logic into SQL code. The code generated consists of control table expressions, ensuring compatibility with most SQL-based data warehouses. This feature significantly reduces manual coding efforts and accelerates data processing.

In conclusion, datasets within the Virtual Data Assistant serve as invaluable assets, providing a holistic view of data and fostering collaborative discussions. With features like live data previews, data transformation templates, and automated SQL code generation, analysts can efficiently explore, analyze, and model data, unlocking valuable insights and making data-driven decisions with ease.

<br>


# Dashboards

Empower Data Analysts with Effortless Dashboard Building

Data visualization is an integral part of data analysis, as it allows users to derive insights and communicate findings effectively. However, creating visually appealing and informative dashboards traditionally involved time-consuming manual processes. With our groundbreaking feature of Natural Language to Dashboard Creation , data analysts can effortlessly build dynamic dashboards using Key Performance Indicators (KPIs) expressed in plain language, all powered by Superset, an open-source data visualization tool.

Below are the steps involved in the dashboard creation process :-

**Select Datasource and Dashboard Description**

{% embed url="<https://www.loom.com/share/eff1fbd56f624f5792a4819c9da3ed2b?sid=e19936c1-9d25-49f0-8ff5-963287b431c9>" %}
Dashboard Creation
{% endembed %}

* The user starts by selecting a data source from which they want to create a dashboard.
* They provide a description or title for the dashboard, outlining its purpose or key focus.

**Add KPIs Using Natural Language**&#x20;

{% embed url="<https://www.loom.com/share/f7006771788f49a796b75f7096eb7f31?sid=0cf320ce-f8f3-45c8-a166-a1c3821679e3>" %}

* Using the Natural Language Dashboard Creation feature, the user adds Key Performance Indicators (KPIs) to the dashboard by expressing their requirements in plain language.
* The system generates SQL queries based on the natural language inputs and displays a preview of the resulting data in real-time on the right-hand side .
* The user has the option to review the generated SQL queries and validate the results presented in the preview section
* If needed, the user can manually modify the SQL queries to fine-tune the data retrieval or apply custom logic for more advanced data processing.
* To further enhance the dashboard's visual appeal and user experience, the user can provide design instructions, such as specifying the chart types, color schemes, layout, and other visual elements.
* An API connects to Superset, the open-source data visualization tool, and adds the specified charts to the dashboard based on the SQL queries and design preferences.

**Further enhance Charts in Superset**&#x20;

Once the charts are added to the dashboard, users have the flexibility to explore and edit the charts further within Superset.

With Superset's rich features, users can perform fine-grained customization, apply additional data filters, and create interactive dashboards tailored to their unique analytical needs.

<br>


# Workbooks

Unlocking the Power of Data Analysis with Interactive Workbooks

Workbooks are dynamic playgrounds designed to empower data analysts, providing them with versatile tools to streamline and enhance their data analytics process. Our feature-rich Workbooks offer a trio of powerful options: Document Book, Query Book, and Expectation Book. Each workbook caters to different aspects of data analysis, offering an intuitive and efficient user experience.

* [**Document Book**](/overview/features/workbooks/document-book)**:** The Document Book revolutionizes the way data analysts work with unstructured data. By harnessing the power of Natural Language Processing (NLP), analysts can effortlessly upload and query documents using simple, everyday language. This interactive capability transforms the analysis of textual data into a seamless and insightful experience, significantly reducing the effort required to derive valuable insights from documents.

* [ **Query Book**](/overview/features/workbooks/query-book)**:** The Query Book empowers data analysts to interact with data through natural language inputs, seamlessly converting plain language queries into SQL. With Text to SQL capabilities, analysts can efficiently and accurately write complex SQL queries without the need for extensive coding knowledge. This feature enhances productivity and enables quick access to the precise data needed for analysis.

* [**Expectation Book**](/overview/features/workbooks/expectations-book)**:** The Expectation Book elevates data quality assurance by allowing analysts to define and execute data quality tasks using expectations. By setting expectations for data, analysts can validate and monitor data accuracy, completeness, and consistency, ensuring reliable and trustworthy analytical outcomes. This feature fosters data confidence and empowers analysts to make data-driven decisions with certainty.

<br>


# Expectations Book

Expectation Books are designed to facilitate data quality assurance and validation. It empowers data analysts to define and run data quality tasks using expectations, ensuring that data meets predefined criteria and adheres to specific requirements. By setting and validating expectations, analysts can instill confidence in the data, leading to more reliable and trustworthy analytical outcomes.

**1. Defining Data Expectations:** In the Expectation Book, data analysts can articulate the expectations they have for the data they are working with. Expectations are predefined rules, conditions, or constraints that data must satisfy to be considered valid and of high quality. For example, an expectation may specify that a certain column should not contain null values or that numeric values should fall within a specific range.

**2. Data Quality Validation:** Once expectations are defined, analysts can execute data quality tasks to validate the data against these expectations. The Expectation Book evaluates the data and highlights any discrepancies or violations of the defined expectations. This validation process helps analysts identify and rectify data anomalies or errors, ensuring the accuracy and completeness of the data.

**3. Customizable Data Expectations:** The Expectation Book offers flexibility in defining expectations based on the specific data and analysis requirements. Analysts can tailor expectations to match the unique characteristics and constraints of their datasets, ensuring that the validation process aligns with their analytical objectives.

**4. Iterative and Data-Driven Decision Making:** The Expectation Book fosters an iterative approach to data analysis. Analysts can repeatedly validate the data against expectations, refining and updating expectations as insights are gained. This iterative process drives data-driven decision-making, as analysts can trust that the data quality is continuously monitored and improved.

**5. Ensuring Data Trustworthiness:** By enforcing data quality through expectations, the Expectation Book helps analysts and stakeholders trust the data they are working with. High-quality data leads to more accurate and reliable insights, reducing the risk of making critical decisions based on flawed or inaccurate information.

**6. Compliance and Governance:** In addition to promoting data trustworthiness, the Expectation Book supports data compliance and governance efforts. By validating data against predefined expectations, organizations can ensure that data adheres to regulatory requirements and internal data standards.

**7. Collaborative Data Quality Assurance:** The Expectation Book facilitates collaboration among data analysts and teams. Analysts can share their defined expectations with colleagues, seek feedback, and jointly establish data quality norms, leading to a collective effort in ensuring data reliability.


# Query Book

Query Book is designed to streamline the process of SQL query generation, documentation, and validation. It empowers data analysts with an intuitive environment to create, store, and preview queries conveniently, eliminating the need to switch between different platforms for exploration and report generation.

{% embed url="<https://www.loom.com/share/9b1704bcfa3949b2ac0f806e36448975?sid=6ae97d33-4682-48f6-adb1-78d85f3f5e1d>" %}
Query Book
{% endembed %}

**1. SQL Query Generation Made Easy:** With the Query Book, data analysts can seamlessly generate SQL queries using natural language inputs. By expressing their data requirements in plain language, analysts can avoid the complexities of writing SQL code manually. This intuitive approach significantly reduces the learning curve and empowers analysts with varying technical backgrounds to create sophisticated queries efficiently.

**2. Documentation and Query Organization:** The Query Book serves as a centralized repository for documenting and organizing SQL queries. Analysts can store and categorize their queries based on different projects, data sources, or analytical needs. This documentation feature ensures that queries are easily accessible and can be revisited or reused in future analyses, promoting consistency and efficiency in data exploration.

**3. Preview Option for Query Validation:** One of the key benefits of the Query Book is the ability to preview queries directly within the platform. Before executing a query on the actual data source, analysts can use the preview option to validate the results and ensure that the query returns the expected data. This real-time validation prevents potential errors and helps analysts fine-tune their queries for accurate data retrieval.

**4. Enhancing Exploratory Data Analysis:** The Query Book proves invaluable for exploratory data analysis. Analysts can iteratively experiment with different queries and explore the data in real-time through previews. This interactive approach facilitates a deeper understanding of the data and empowers analysts to make data-driven decisions based on a comprehensive view of the dataset.

**5. Efficient Report Generation:** By enabling the creation of documented and validated queries, the Query Book expedites report generation. Analysts can quickly refer back to previously saved queries and use them as building blocks for report components. This feature streamlines the reporting process and ensures consistency in the data used across various reports.

**6. Collaboration and Knowledge Sharing:** The Query Book also fosters collaboration among team members. Analysts can share their documented queries with colleagues, promoting knowledge sharing and facilitating peer review. This collaborative aspect enhances the quality of analyses and enables the collective growth of the analytics team.


# Document Book

Designed to streamline the process of analyzing textual documents and extracting

The Document Workbook is a cutting-edge feature designed to streamline the process of analyzing textual documents and extracting valuable insights without the need for manual parsing, crawling, or extensive reading. This innovative tool empowers analysts, researchers, and knowledge workers to efficiently upload documents, ask targeted questions related to the content, and receive precise answers within a scoped and contextually relevant workspace.

Key Features and Benefits:

1. **Seamless Document Upload**: The Document Workbook enables users to effortlessly upload textual documents in PDF format This eliminates the time-consuming task of manually extracting information from documents, allowing analysts to dive straight into their analysis.
2. **Scoped Workbooks for Focused Analysis**: Each uploaded document is associated with a scoped workbook, creating a segregated environment where questions can be asked exclusively about the content of that specific document. This focus ensures that queries and answers remain relevant to the material being examined.
3. **Contextual Questioning**: Analysts can ask questions related to the uploaded documents using natural language queries. The Document Workbook's advanced natural language processing capabilities enable it to understand the context and extract the most pertinent information from the documents, significantly reducing the need for manual intervention.
4. **Precise Answers in Real-Time**: By leveraging powerful AI algorithms, the Document Workbook provides real-time responses to questions asked by analysts. This instant access to targeted information enhances efficiency and facilitates faster decision-making.
5. **Intelligent Data Indexing**: The Document Workbook automatically indexes the uploaded documents, allowing for rapid search and retrieval of relevant passages. This feature proves invaluable when navigating large volumes of documentation.
6. **Collaborative Workspace**: The Document Workbook fosters collaboration by allowing multiple analysts to work together within the same scoped environment. This ensures that insights are shared and discussed efficiently, facilitating a synergistic approach to document analysis.
7. **Audit Trail and Version Control**: The system maintains a comprehensive audit trail of questions asked, responses received, and changes made within the workbook. Version control capabilities help track document updates, ensuring data accuracy and preserving historical context.
8. **User-Friendly Interface**: With an intuitive and user-friendly interface, the Document Workbook requires minimal training, making it accessible to both technical and non-technical users. Its simplicity enhances user adoption and efficiency.
9. **Security and Data Privacy**: The Document Workbook prioritizes data security and privacy, with robust encryption and access controls to safeguard sensitive information. Compliance with industry standards ensures that confidential data remains protected.
10. **Customizable Workflows**: Tailored to the specific needs of analysts and researchers, the Document Workbook allows for customizable workflows and integration with existing document management systems.

By leveraging the Document Workbook, analysts can now gain valuable insights from textual documents without having to sift through extensive pages of documentation manually. The feature's ability to scope workbooks, ask contextually relevant questions, and provide precise answers in real-time transforms document analysis into a seamless and efficient process, driving informed decision-making and enhanced collaboration among teams. Embrace the Document Workbook and revolutionize the way you interact with textual data, making knowledge discovery and analysis faster, more accurate, and more impactful.

Below Video

Document&#x20;


# API Documentation

Below Graph QL API would help applications integrate with external applications.&#x20;

Request Structure

```
mutation QueryWorkBook($queryWorkBookId: String!, $question: String!) {
  queryWorkBook(id: $queryWorkBookId, question: $question) {
    id
    message
    documentBookId
    creatorId
    createdAt
    updatedAt
    creator {
      id
    }
  }
}
```

Response Structure

Sample Response&#x20;

```json5
{"data":{"queryWorkBook":[{"id":"136","message":"Hello ","documentBookId":"26","creatorId":"4baba2fb-d0f8-4678-9d11-51ae1f7843da","createdAt":1690452915701,"updatedAt":1690452915703,"creator":{"id":"4baba2fb-d0f8-4678-9d11-51ae1f7843da"}},{"id":"137","message":" Hi there! How can I help you?","documentBookId":"26","creatorId":"system","createdAt":1690452918741,"updatedAt":1690452918742,"creator":{"id":"system"}}]}}





```


# VDA in Docker

Docker Installation Guide for different OS(s)

{% embed url="<https://docs.docker.com/engine/install/>" %}

### Windows

Install Docker Desktop&#x20;

{% embed url="<https://docs.docker.com/desktop/install/windows-install/>" %}

{% embed url="<https://desktop.docker.com/win/main/amd64/Docker%20Desktop%20Installer.exe?_gl=1*iak05f*_ga*NDE1Mzk2OTkyLjE3MTc2Njc0OTk.*_ga_XJWPQMJYHQ*MTcxNzY2NzQ5OS4xLjEuMTcxNzY2NzYzMi41OS4wLjA>." %}
Link to Download
{% endembed %}

Run following command on Power shell

`wsl --set-default ubuntu`&#x20;

To check docker information run following command on Power shell

docker info&#x20;

### Ubuntu

Set up Docker's `apt` repository.

**Add Docker's official GPG key**:

`sudo apt-get update sudo apt-get install ca-certificates curl sudo install -m 0755 -d /etc/apt/keyrings sudo curl -fsSL https://download.docker.com/linux/ubuntu/gpg -o /etc/apt/keyrings/docker.asc sudo chmod a+r /etc/apt/keyrings/docker.asc`

**Add the repository to Apt sources:**

`echo`\
`"deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.asc] https://download.docker.com/linux/ubuntu`\
`$(. /etc/os-release && echo "$VERSION_CODENAME") stable" |`\
`sudo tee /etc/apt/sources.list.d/docker.list > /dev/null sudo apt-get update`

**To install the latest version, run**:

`sudo apt-get install docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin`

Verify that the Docker Engine installation is successful by running the `hello-world` image.

`sudo docker run hello-world`

Post Installation steps to run docker without root user

Create the `docker` group.

`sudo groupadd docker`

Add your user to the `docker` group.

`sudo usermod -aG docker $USER`

Log out and log back in so that your group membership is re-evaluated.

`newgrp docker`

Verify that you can run `docker` commands without `sudo`.

docker run hello-world

### Install git on OS

{% embed url="<https://git-scm.com/download/win>" %}

### System Requirements

Minimum requirements allocated to Docker should be&#x20;

| RAM   | CPU | Disk  |
| ----- | --- | ----- |
| 12 GB | 4   | 20 GB |

To change the memory allocation for Docker, please visit `Preferences -> Resources -> Advanced` in your Docker Desktop.

### Procedure

#### Clone VDA Installer Git Repository

```
git clone https://github.com/microui-team/vdainstaller.git
```

#### Setting Environment Variables&#x20;

Change the values for parameters as per your keys .&#x20;

```
// Some code
# Registry from which we are pulling the image
IMAGE_REGISTRY=microuidigital
GIT_SHA=latest

# Database configurations
POSTGRES_PASSWORD=postgrespwd  # Change the database password if desired.
DATABASE_URL=postgresql://postgres:postgrespwd@postgres:5432/postgres?schema=public

# Keycloak
KEYCLOAK_URL=http://localhost:8080
KEYCLOAK_ADMIN_USERNAME=whiteklay
KEYCLOAK_ADMIN_PASSWORD=password

# Superset engine specific environment variables - update based on your configuration
ORIG_SUPERSET_URL=http://localhost:8088
SUPERSET_USER=admin
SUPERSET_PASSWORD=admin
ORIG_SUPERSET_TOKEN=123456abc

# AWS credentials
AWS_S3_BUCKET_NAME={your-s3-bucket-name}
AWS_ACCESS_KEY_ID={your-aws-access-key-id}
AWS_SECRET_ACCESS_KEY={your-aws-secret-access-key}

# Realm name
REALM_NAME=ssm-tool

# Environment variables for ingestion
DATA_INGESTION_URL=http://{server-ip}:8083
GIT_TOKEN={your-git-token}
GIT_REPONAME=vda-notebook

# Message queue configurations
MESSAGE_RESPONSE_QUEUE_NAME=amundsen_response_queue
MESSAGE_QUEUE_NAME=amundsen_queue
MESSAGE_QUEUE_PORT=5672
MESSAGE_QUEUE_USER=whiteklay
MESSAGE_QUEUE_PASSWORD=password
MESSAGE_QUEUE_HOST=rabbitmq

# API keys
OPENAI_API_KEY={your-openai-api-key}
HUGGINGFACEHUB_API_TOKEN={your-huggingface-api-token}

```

#### Start Docker Services

&#x20;

```
cd vdainstaller
sh docker-compose-up.sh
```

The docker compose up script calls the script which can also be used directly&#x20;

`docker compose -f vda-deploy.yaml --env-file vda/.env-vda up -d`

#### Stopping the Services

```
cd vdainstaller
sh docker-compose-down.sh
```

The docker compose up script calls the script which can also be used directly&#x20;

`docker compose -f vda-deploy.yaml --env-file vda/.env-vda down`

#### Cleaning up the environment

####

```
cd vdainstaller
sh docker-volume-prune.sh
```

The docker compose up script calls the script which can also be used directly&#x20;

`docker compose -f vda-deploy.yml --env-file vda/.env-vda down`&#x20;

`docker volume prune -a`


# Data Catalog

## What is a Data Catalog

A data catalog is a comprehensive inventory of data assets within an organization, designed to help users find, understand, and use data effectively.

## Metadata in Data Catalog

Metadata is the descriptive information about data, providing context and meaning. In a data catalog, metadata can be categorized into several types:

1. **Technical Metadata**:
   * **Schema Information**: Details about the structure of the data, such as tables, columns, data types, indexes, and constraints.
   * **Data Source Information**: Information about where the data originates, such as database names, server locations, and connection strings.
   * **Relationship**: Defines how a particular data is related to other data sets .
2. **Business Metadata**:
   * **Business Glossary**: Definitions and descriptions of business terms and concepts to ensure a common understanding across the organization.
   * **Data Ownership**: Information about who is responsible for the data, including data stewards and data owners.
   * **Usage Context**: Information about how and why the data is used in business processes and decision-making.
3. **Operational Metadata**:
   * **Data Quality Metrics**: Information about the accuracy, completeness, consistency, and timeliness of the data.
   * **Access and Usage Statistics**: Details about who accessed the data, when it was accessed, and how frequently it is used.
   * **Processing Metadata**: Information about data processing jobs, such as ETL (Extract, Transform, Load) processes, including job schedules, statuses, and logs.
4. **Governance Metadata**:
   * **Policies and Compliance**: Information about data governance policies, regulatory requirements, and compliance statuses.
   * **Security Information**: Details about data security measures, including encryption, masking, and access controls.

### VDA Data Catalog :

Data Catalog in VDA consists of below components&#x20;

### Data Source

A data source is a logical or physical grouping of related data assets, which can include tables, files, objects, or even dashboards

{% content-ref url="/pages/pur84mDh9mPIooI2RTa4" %}
[DataSource](/how-to-guides/data-catalog/datasource)
{% endcontent-ref %}

### Datasets

A dataset is a collection of related data, typically organized into tables, files, or objects, that is used for analysis, reporting, or other data-driven tasks

{% content-ref url="/pages/FtpXlwGCWg0JQmzpqKvv" %}
[Datasets](/how-to-guides/data-catalog/datasets)
{% endcontent-ref %}


# DataSource

## What is a Data Source

A Datasource is an entity within the Virtual Data Assistant that serves as a container for a collection of metadata. Metadata refers to information about datasets, such as data source location, schema, data types, and other relevant properties. Essentially, a Datasource is like a virtual folder that groups related datasets together, making it easier for users to manage and access data efficiently.

## How to Create a new Data Source

Navigate to Datasources and click create to create a new data source

<figure><img src="https://4133454587-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FfOPtN4DPW3XVXKtINi7z%2Fuploads%2F2Q4zcKjCb6QWgNZLfp6J%2Fcreatedatasource.png?alt=media&#x26;token=44688e48-2fa1-4dfc-bdc5-7854a9499b22" alt=""><figcaption><p>create new data source</p></figcaption></figure>

Select the connector( source type ) radio button as per the required data source

<figure><img src="https://4133454587-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FfOPtN4DPW3XVXKtINi7z%2Fuploads%2F1Guc77yQEtdaPGOg6VYR%2FScreen%20Shot%202023-07-29%20at%205.55.49%20PM.png?alt=media&#x26;token=6e4e9b96-511d-4876-a8ab-eaa1de3d2a11" alt=""><figcaption><p>Create a datasource</p></figcaption></figure>

#### Advance Properties

If you do not want all tables, then you can add table names in Include or Exclude boxes

**Include**: Fetch only for these table(s)

**Exclude**: Fetch all excluding mentioned table(s)

Connectors are modules or plugins that establish connections to specific data sources or databases. VDA offers multiple connectors to popular datasources e.g PostgreSQL, MySQL etc. When users select a connector and provide the necessary connection details, it establishes a link to the data source, allowing access to the data within that source.

### Connector Parameters

#### Postgres

<figure><img src="https://4133454587-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FfOPtN4DPW3XVXKtINi7z%2Fuploads%2FGkNXWcWeF89KQMldBF58%2FScreen%20Shot%202024-06-24%20at%204.55.23%20PM.png?alt=media&#x26;token=02de7f6f-4de9-4dfe-a1ac-9f05934072dc" alt=""><figcaption><p>Postgres Connection</p></figcaption></figure>

Postgres data source type allows below two methods :-

* **Detail**:\
  Select this parameter when you are connecting using standard username , password and port .
* **URL**:\
  Select this parameter when you have a custom URL connection string Below is an example of a connection string

`postgresql://user:password@hostname:port/database`

#### Oracle

<figure><img src="https://4133454587-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FfOPtN4DPW3XVXKtINi7z%2Fuploads%2FeyqC5BUtR8wijiRFzkFo%2Fimage.png?alt=media&#x26;token=6d5eaf59-13e3-4145-a8b5-da452814bc81" alt=""><figcaption><p>Oracle Connection</p></figcaption></figure>

* **Description**: Use this method when connecting to an Oracle database using standard credentials and connection parameters.
* **Parameters**:
  * **Role**: The role assigned to the user (e.g., DBA, Developer).
  * **Host**: The hostname or IP address of the Oracle server.
  * **User**: The username for the Oracle database.
  * **Password**: The password for the Oracle database.
  * **Schema**: The schema within the Oracle database to which the user has access.
  * **Service**: The Oracle service name (e.g., ORCL).

<figure><img src="https://4133454587-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FfOPtN4DPW3XVXKtINi7z%2Fuploads%2FY5DpaxwifjBttlz4pKKy%2Fimage.png?alt=media&#x26;token=852c7aca-df07-4851-bca9-01d5169f1837" alt=""><figcaption><p>Oracle Connection - Advance</p></figcaption></figure>

**URL Method**:

* **Description**: Use this method when you have a custom URL connection string for the Oracle database.

`jdbc:oracle:thin:@//oracle.example.com:1521/ORCL`

#### Kafka

<figure><img src="https://4133454587-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FfOPtN4DPW3XVXKtINi7z%2Fuploads%2FFHLZ7C8YexIBsGr4rg3i%2Fimage.png?alt=media&#x26;token=6c179ae1-1796-48b2-a12b-20ba15212420" alt=""><figcaption><p>Kafka Connection - Advance</p></figcaption></figure>

**Name**:

* **Description**: The name of the Kafka connection.
* **Example**: `Kafka_Prod_Cluster`

**URL** :

* **Description**: The base URL for the Kafka service.
* **Importance**: Provides the endpoint for connecting to the Kafka service.
* **Example**: `kafka.example.com`

**Broker URL**:

* **Description**: A comma-separated list of host and port pairs that are the addresses of the Kafka brokers.
* **Importance**: Specifies the Kafka brokers to connect to.
* **Example**: `kafka1.example.com:9092,kafka2.example.com:9092`

**User**:

* **Description**: The username for the Kafka connection.
* **Importance**: Used for authenticating the user accessing the Kafka cluster.
* **Example**: `kafka_user`

**Password**:

* **Description**: The password for the Kafka connection.
* **Importance**: Secures the connection by authenticating the user.
* **Example**: `securepassword`

**Secure Connection**:

* **Description**: A checkbox option to enable a secure connection.
* **Importance**: Ensures that the data transmitted between the client and Kafka brokers is encrypted.
* **Example**: `Checked` or `Unchecked`

**Included Tables**:

* **Description**: A list of specific tables/topics to be included in the data ingestion process.
* **Importance**: Allows for targeted data ingestion, focusing on relevant tables/topics.
* **Example**: `topic1, topic2, topic3`

**Excluded Tables**:

* **Description**: A list of specific tables/topics to be excluded from the data ingestion process.
* **Importance**: Prevents unnecessary or irrelevant data from being ingested.
* **Example**: `topic4, topic5`

#### Hive

<figure><img src="https://4133454587-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FfOPtN4DPW3XVXKtINi7z%2Fuploads%2FIGuATRbZSa5tACjoO7F4%2Fimage.png?alt=media&#x26;token=7482cbef-619a-4501-b955-54e13767ba22" alt=""><figcaption><p>Hive Connection - Advance</p></figcaption></figure>

**Detail Method**

**Name**:

* **Description**: The name of the Hive connection.
* **Importance**: Identifies the specific Hive data source within the VDA
* **Example**: `Hive_Prod_Cluster`

**Metastore URL**:

* **Description**: The URL of the Hive Metastore service.
* **Importance**: Provides the endpoint for connecting to the Hive Metastore, which manages metadata for Hive tables and databases.
* **Example**: `thrift://metastore.example.com:9083`

**Hive URL**:

* **Description**: The URL of the Hive server.
* **Importance**: Provides the endpoint for connecting to the Hive server for executing queries and accessing data.
* **Example**: `jdbc:hive2://hive.example.com:10000/default`

**Database**:

* **Description**: The name of the specific database within the Hive server to connect to.
* **Importance**: Specifies the target database for data operations.
* **Example**: `default`

**Included Tables**:

* **Description**: A list of specific tables to be included in the data ingestion process.
* **Importance**: Allows for targeted data ingestion, focusing on relevant tables.
* **Example**: `table1, table2, table3`

**Excluded Tables**:

* **Description**: A list of specific tables to be excluded from the data ingestion process.
* **Importance**: Prevents unnecessary or irrelevant data from being ingested.
* **Example**: `table4, table5`

#### SQL Server

<figure><img src="https://4133454587-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FfOPtN4DPW3XVXKtINi7z%2Fuploads%2FKvvaPuzmfK0Vrh0PrN5w%2Fimage.png?alt=media&#x26;token=ab0dca70-bdc3-45f8-be73-f8a359b4dd9a" alt=""><figcaption><p>SQL Server Connection</p></figcaption></figure>

**Detail Method**

**Name**:

* **Description**: The name of the SQL Server connection.
* **Importance**: Identifies the specific SQL Server data source within the VDA
* **Example**: `SQLServer_Prod`

**Host**:

* **Description**: The hostname or IP address of the SQL Server.
* **Importance**: Specifies the server where the SQL Server is hosted.
* **Example**: `sqlserver.example.com`

**Port**:

* **Description**: The port number used to connect to the SQL Server.
* **Importance**: Specifies the network port for the SQL Server connection.
* **Example**: `1433`

**User**:

* **Description**: The username for the SQL Server database.
* **Importance**: Used for authenticating the user accessing the SQL Server.
* **Example**: `db_user`

**Password**:

* **Description**: The password for the SQL Server database.
* **Importance**: Secures the connection by authenticating the user.
* **Example**: `securepassword`

**Database**:

* **Description**: The name of the specific database within the SQL Server to connect to.
* **Importance**: Specifies the target database for data operations.
* **Example**: `mydatabase`

**Schema**:

* **Description**: The schema within the SQL Server database.
* **Importance**: Defines the organizational structure of tables within the database.
* **Example**: `dbo`

**Included Tables**:

* **Description**: A list of specific tables to be included in the data ingestion process.
* **Importance**: Allows for targeted data ingestion, focusing on relevant tables.
* **Example**: `table1, table2, table3`

**Excluded Tables**:

* **Description**: A list of specific tables to be excluded from the data ingestion process.
* **Importance**: Prevents unnecessary or irrelevant data from being ingested.
* **Example**: `table4, table5`

**URL Method**

**Description**: Use this method when you have a custom URL connection string for SQL Server, incorporating all necessary connection details.

`jdbc:sqlserver://sqlserver.example.com:1433;databaseName=mydatabase;user=db_user;password=securepassword;schema=dbo`

#### My SQL

<figure><img src="https://4133454587-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FfOPtN4DPW3XVXKtINi7z%2Fuploads%2F2MBfKbZ4bbsXbCACzBUh%2Fimage.png?alt=media&#x26;token=363b1c25-02c0-4242-86e4-7cfc46b5aaf2" alt=""><figcaption></figcaption></figure>

**Detail Method**

**Name**:

* **Description**: The name of the MySQL connection.
* **Importance**: Identifies the specific MySQL data source within the VDA
* **Example**: `MySQL_Prod`

**Host**:

* **Description**: The hostname or IP address of the MySQL server.
* **Importance**: Specifies the server where the MySQL database is hosted.
* **Example**: `mysql.example.com`

**Port**:

* **Description**: The port number used to connect to the MySQL server.
* **Importance**: Specifies the network port for the MySQL connection.
* **Example**: `3306`

**User**:

* **Description**: The username for the MySQL database.
* **Importance**: Used for authenticating the user accessing the MySQL database.
* **Example**: `db_user`

**Password**:

* **Description**: The password for the MySQL database.
* **Importance**: Secures the connection by authenticating the user.
* **Example**: `securepassword`

**Database**:

* **Description**: The name of the specific database within the MySQL server to connect to.
* **Importance**: Specifies the target database for data operations.
* **Example**: `mydatabase`

**Schema**:

* **Description**: The schema within the MySQL database.
* **Importance**: Defines the organizational structure of tables within the database.
* **Example**: `public`

**Included Tables**:

* **Description**: A list of specific tables to be included in the data ingestion process.
* **Importance**: Allows for targeted data ingestion, focusing on relevant tables.
* **Example**: `table1, table2, table3`

**Excluded Tables**:

* **Description**: A list of specific tables to be excluded from the data ingestion process.
* **Importance**: Prevents unnecessary or irrelevant data from being ingested.
* **Example**: `table4, table5`

**URL Method**

**Description**: Use this method when you have a custom URL connection string for MySQL, incorporating all necessary connection details.

`jdbc:mysql://mysql.example.com:3306/mydatabase?user=db_user&password=securepassword`

**List of Available Connectors (Grouped by Category)**

**Database Connectors (54)**\
Amazon Athena, Azure SQL, BigQuery, BigTable, BurstIQ, Cassandra, Clickhouse, Cockroach, Couchbase, Custom Database, Databricks, Datalake, Db2, DeltaLake, DomoDatabase, Doris, Dremio, Druid, DynamoDB, Epic, Exasol, Glue, Greenplum, Hive, Iceberg, Impala, Informix, MariaDB, Microsoft Access, Microsoft Fabric, MongoDB, MSSQL, MySQL, Oracle, PinotDB, Postgres, Presto, Redshift, SAS, SQLite, SSAS, Salesforce, SAP ERP, SAP HANA, ServiceNow, SingleStore, Snowflake, StarRocks, Synapse, Teradata, Timescale, Trino, Unity Catalog, Vertica

**Dashboard Connectors (20)**\
Custom Dashboard, Domo Dashboard, Grafana, Hex, Lightdash, Looker, Metabase, MicroStrategy, Mode, PowerBI, PowerBI Report Server, QlikCloud, QlikSense, QuickSight, Redash, Sigma, Ssrs, Superset, Tableau, ThoughtSpot

**Pipeline Connectors (25)**\
Airbyte, Airflow, Backend, Custom Pipeline, DBT Cloud, Dagster, Data Factory, Databricks Pipeline, Domo Pipeline, Fivetran, Flink, Glue Pipeline, Kafka Connect, Kinesis Firehose, Matillion, Microsoft Fabric Pipeline, Mulesoft, Nifi, OpenLineage, SSIS, Snowplow, Spark, Spline, Stitch, Wherescape

**Messaging Connectors (5)**\
Custom Messaging, Kafka, Kinesis, Pulsar, Redpanda

**LLM (Large Language Model) Connectors (8)**\
Anthropic, Azure OpenAI, Bedrock, Custom LLM, HuggingFace, Ollama, OpenAI, VertexAI

**Storage Connectors (4)**\
Adls, Custom Storage, GCS, S3

**Search Connectors (3)**\
Custom Search, Elasticsearch, OpenSearch

**Drive Connectors (4)**\
Custom Drive, Google Drive, SFTP, SharePoint

**ML Model Connectors (5)**\
Custom ML Model, Mlflow, SageMaker, Sklearn, VertexAI

**Metadata Connectors (7)**\
Alation, AlationSink, Amundsen, Atlas, Collibra, MetadataES, OpenMetadata

**API & Security Connectors (2)**\
REST API, Ranger (Security)


# Datasets

Datasets are the lifeblood of data analysis, and   Virtual Data Assistant (VDA) empowers users with a feature-rich environment to harness the full potential of data. Datasets within the VDA serve as powerful entities that showcase data as a product, providing valuable insights and enabling data-driven decision-making. Let's delve into the key features of Datasets:

Types of Datasets

* Live

&#x20;Live datasets refer to datasets that are directly connected to real  data sources. These datasets are  updated and reflect the most current information available from the connected data sources. VDA establishes and maintains live connections to these data sources, ensuring that any changes or updates in  metadata of source data are immediately reflected in the corresponding dataset.

* Custom

Custom datasets, on the other hand, are datasets that are created and defined by users to meet specific analytical requirements. Unlike live datasets, custom datasets are not directly linked to real-time data sources. Instead, users define the data they want to include in the dataset and may apply data transformations to manipulate and refine the data to suit their analytical needs.

<br>

**1. Metadata, Overview, and Tags:**&#x20;

{% embed url="<https://www.loom.com/share/134cf121f45f4ee29d6524488a6ea07b>" %}

Datasets are accompanied by comprehensive metadata, offering essential information about the data they contain. This metadata includes a list of columns, descriptions for each column, and data types, providing users with a clear understanding of the dataset's structure and content. Additionally, an Overview section offers a concise summary of the dataset's key characteristics and significance.

To further enhance organization and accessibility, users can add tags to datasets. Tags serve as labels or keywords, making it easier to search and categorize datasets effectively.

**2. Preview - See Live Data**

{% embed url="<https://www.loom.com/share/6932d5f8efcd48eeb665fe70fe2fe446?sid=5a3f68ce-7063-4fe1-8236-4cd1c5d510ae>" %}

Datasets come to life with the "Preview" feature, offering users a glimpse of the actual data contained within. This live data preview allows analysts to quickly assess the dataset's contents before diving into in-depth analyses. It also aids in identifying any potential data issues or anomalies that may require attention.

**3. Stories - Fostering Collaborative Discussions:**&#x20;

{% embed url="<https://www.loom.com/share/466bff30b64841a597b4c103dd5b66da?sid=6df492dd-23cd-40f1-8e8e-b8b719764fdb>" %}

The "Stories" section provides a collaborative platform for users to engage in discussions about datasets. It serves as a dynamic forum where data analysts, stakeholders, and team members can share insights, ask questions, and exchange valuable perspectives related to the dataset. Additionally, users can upload important documents, transforming the Stories section into a central hub for discussions and knowledge-sharing.

**4. Transformations - Empowering Custom Datasets**

{% embed url="<https://www.loom.com/share/311f7a2b249a4d17b09bf2474521f408?sid=a1ab490a-9cac-45ba-a767-e8c653753e5b>" %}
Create Transformations
{% endembed %}

For custom datasets, the VDA offers the "Transformations" feature. This powerful tool enables users to define data transformations using templates, facilitating data modeling and customization. Analysts can upload transformation logic and document the process, streamlining the preparation and refinement of custom datasets.

**5. Publish and Generate SQL - Effortless Code Generation:**&#x20;

{% embed url="<https://www.loom.com/share/2086acbcc07d43f0b6029b44c5cdb7a7?sid=a73eef3f-0983-44c1-955c-8db73bffd648>" %}

Users can publish versions for transformations and also utilize the generate SQL option to generate code for the defined transformations .

The "Generate SQL" feature leverages the power of OpenAI to automate code generation based on the logic defined in transformations. With just a few clicks, users can effortlessly convert their transformation logic into SQL code. The code generated consists of control table expressions, ensuring compatibility with most SQL-based data warehouses. This feature significantly reduces manual coding efforts and accelerates data processing.

#### 6. ER Diagram Tab

The **ER Diagram** tab visualizes the entity-relationship model of the dataset, showing tables, columns, and relationships. This graphical representation helps users understand the data schema and how different entities are interconnected. It highlights primary and foreign key relationships, making it easier to comprehend complex data structures. Supports database design and optimization by revealing the dataset's architecture. Useful for both technical and non-technical users to gain a high-level overview of data organization.

In conclusion, datasets within the Virtual Data Assistant serve as invaluable assets, providing a holistic view of data and fostering collaborative discussions. With features like live data previews, data transformation templates, and automated SQL code generation, analysts can efficiently explore, analyze, and model data, unlocking valuable insights and making data-driven decisions with ease.

<br>


# Exploration

VDA offers powerful and comprehensive data exploration capabilities, allowing users to efficiently locate and access the data they need. This functionality supports searches based on various criteria, including table names, column names, database names, schema names, and tags associated with tables. Here’s a detailed overview of these search capabilities:

#### Table Search

**Description**: Allows users to search for data by specifying the table name.

* **Use Case**: Ideal for users who know the specific table they need to access or analyze.
* **Example**: A user can quickly locate the "Bank" table to review the data
* **Benefit**: Streamlines the process of finding relevant tables, saving time and effort.

### Steps

Click on Dataset&#x20;

<figure><img src="/files/HsnKr7VQaUy9txCZx8sq" alt=""><figcaption><p>Data Exploration</p></figcaption></figure>

Enter the table name followed by \*(if you don't remember exact name)

<figure><img src="/files/wY7wkVfpwPyCOtI6AYkK" alt=""><figcaption><p>Table Search</p></figcaption></figure>

#### Column Search

**Description**: Enables users to search within tables based on column names.

* **Use Case**: Useful when users are looking for specific data attributes or fields across multiple tables.
* **Example**: A user searching for "card" can find all tables containing this column.
* **Benefit**: Facilitates detailed and targeted data searches, enhancing data discovery.

<figure><img src="/files/Ecpl8Es8n7WjdvufCVHH" alt=""><figcaption><p>Column Search</p></figcaption></figure>

#### Database Search

**Description**: Supports searches by database names(Ex. postgres, Oracle etc), helping users locate data within specific databases.

* **Use Case**: Essential for organizations with multiple databases, allowing users to focus their search on a particular database.
* **Example**: A user can search for the "postgres" to find all relevant tables and datasets within this type of database.
* **Benefit**: Simplifies navigation through large, complex data environments by narrowing down search scope.

<figure><img src="/files/RkPYocQkieNYYj4aNe3w" alt=""><figcaption><p>Database Search</p></figcaption></figure>

#### Schema Search

**Description**: Allows users to search based on schema names within a database.

* **Use Case**: Beneficial for users who need to access data organized under specific schemas.
* **Example**: Searching for the "public" schema will return all tables and objects within that schema.
* **Benefit**: Provides an organized view of data structures, making it easier to locate data assets.

<figure><img src="/files/m3z3l6sKFzZ8T6EJVT5X" alt=""><figcaption><p>Schema Search</p></figcaption></figure>

#### Tag-Based Search

**Description**: Enables searches using tags associated with tables, allowing for thematic or categorical searches.

* **Use Case**: Useful for users who categorize data tables with tags such as "financial," "archived," or "sensitive."
* **Example**: A user searching for the tag "sensitive" will find all tables marked as containing sensitive information.
* **Benefit**: Enhances data organization and discoverability, enabling users to find data based on specific characteristics or purposes.

<figure><img src="/files/80bjlMijRNe5IGzpkJ5O" alt=""><figcaption><p>Tag Search</p></figcaption></figure>

<figure><img src="/files/nIOHuOTxL1R2Qr1CSpll" alt=""><figcaption><p>Tag Search</p></figcaption></figure>

#### Data Exploration Benefits

1. **Efficiency**: The ability to search based on table, column, database, schema, and tags significantly reduces the time spent locating specific data.
2. **Comprehensiveness**: By supporting multiple search criteria, the VDA tool ensures that users can find exactly what they need, regardless of their starting point or the detail they have.
3. **Usability**: A user-friendly search interface allows both technical and non-technical users to perform searches with ease, promoting wider adoption and usage.
4. **Data Discovery**: Enhanced search capabilities foster better data discovery, enabling users to uncover hidden insights and relationships within the data.
5. **Organization**: Helps maintain a well-organized data environment, where data assets are easily accessible and efficiently categorized.


# Data Quality

Data quality refers to the assessment of a dataset's overall state. It measures objective elements such as completeness, accuracy, and consistency. But it also measures more subjective factors, such as how well-suited a dataset is to a particular task.

Data Quality within VDA is addressed using the below features

### Profiling

Data profiling is the process of uncovering and investigating data quality issues, such as duplication, inconsistency, inaccuracies, and incompleteness. It involves analyzing one or multiple data sources and gathering metadata that reflects the condition of the data. This metadata allows data stewards to trace and investigate the origins of data errors effectively.

{% content-ref url="/pages/Q3g4WXEofv2VjNhXPKIy" %}
[Profiling](/how-to-guides/data-quality/profiling)
{% endcontent-ref %}

### Expectations

Expectations refer to predefined criteria or rules that data must meet to be considered accurate, complete, consistent, and reliable. These expectations help ensure data integrity and usability by setting standards for data quality.

{% content-ref url="/pages/wgGPdIP4Rp2iCaQSy7Eh" %}
[Expectations](/how-to-guides/data-quality/expectations)
{% endcontent-ref %}


# Expectations

VDA's Expectation Books are designed to facilitate data quality assurance and validation. It empowers data analysts to define and run data quality tasks using expectations, ensuring that data meets predefined criteria and adheres to specific requirements. By setting and validating expectations, analysts can instill confidence in the data, leading to more reliable and trustworthy analytical outcomes.

### How to Create Expectations

Expectations are part of Workbooks within VDA  . Navigate to workbooks tab and select "Expectation Book" .&#x20;

<figure><img src="/files/rjAXWREDH1qvmZbyg239" alt=""><figcaption><p>Expecatation Workbook</p></figcaption></figure>

Once selected provide a name and description for the expectation book and select a Datasource on which the expectation needs to run .A Expectation book is a collection of data expectations which can be grouped as per datasources or business use-cases .

<figure><img src="/files/pxBVJynASHD0XLn314P6" alt=""><figcaption><p>Create a Expectation</p></figcaption></figure>

Once a book is created click on the + button on the right bottom of the screen  to create an expectation .

<figure><img src="/files/G7uUoIT5IuwBBxr1DHqe" alt=""><figcaption><p>Expectation Detail Page</p></figcaption></figure>

Select Templated or Custom

<figure><img src="/files/4sqXmHREcS6obCt1R8Jh" alt=""><figcaption><p>Create Expectation Type</p></figcaption></figure>

Expectations are of two types

* **Templated**&#x20;

Templated expectations are predefined criteria provided by VDA, commonly used across the data industry. These expectations cover a wide range of standard data quality checks and are frequently updated. They are stored in a public repository that auto-syncs, ensuring users always have access to the latest set of expectations.  Details for [templated  expectations ](/how-to-guides/data-quality/expectations) can be found in the next sections .&#x20;

&#x20;

* **Custom**

Custom expectations, on the other hand, are criteria created by organizations to meet specific needs when no relevant templated expectation is available. These allow organizations to define their own data quality rules tailored to their unique data requirements and quality standards.Templated expectations are created using templates once selected you will be redirected to create a expectation template .

If you have selected Templated, then from bottom panel select Template name

Enter the details like Database, Table, column(s) and expected values

<figure><img src="/files/d1CFF1wmIh4wQCny3NDe" alt=""><figcaption><p>Create Expectation for Template</p></figcaption></figure>

Expectations form will change based on the template below is an example for a template for a expectation which expects column as not null .

<figure><img src="/files/gBQY01PVgorrWXSKiEg5" alt=""><figcaption></figcaption></figure>

Enter the details like Database, Table, column(s) and expected values

<figure><img src="/files/ymjpjLyHhrTjI74q0RFW" alt=""><figcaption></figcaption></figure>

Expectation will get created

Now you can Preview your Expectation

####


# Templated Expectations

VDA provides a growing repository of out-of-the-box expectations so that users can benefit without the need to create and define their own. These predefined criteria cover a wide range of standard data quality checks and are categorized at both the table and column levels.

**Table-Level Expectations:**

* Table Row Count to Equal
* Table Row Count to be Between
* Table Column Count to Equal
* Table Column Count to be Between
* Table Column Name to Exist
* Table Column to Match Set
* Table Custom SQL Test
* Table Row Inserted Count to be Between
* Expect table columns to match ordered list
* Expect table row count to be between
* Expect table row count to equal

**Column-Level Expectations:**

* **expect\_column\_max\_to\_be\_between**: Expect the column maximum to be between a minimum and a maximum value.
* **expect\_column\_mean\_to\_be\_between**: Expect the column mean to be between a minimum and a maximum value (inclusive).
* **expect\_column\_median\_to\_be\_between**: Expect the column median to be between a minimum and a maximum value.
* **expect\_column\_min\_to\_be\_between**: Expect the column minimum to be between a minimum value and a maximum value.
* **expect\_column\_values\_to\_be\_in\_set**: Expect each column value to be in a given set.
* **expect\_column\_values\_to\_be\_in\_type\_list**: Expect a column to contain values from a specified type list.
* **expect\_column\_values\_to\_be\_null**: Expect the column values to be null.
* **expect\_column\_values\_to\_be\_of\_type**: Expect a column to contain values of a specified data type.
* **expect\_column\_values\_to\_be\_unique**: Expect each column value to be unique.
* **expect\_column\_values\_to\_not\_be\_null**: Expect the column values to not be null.

These templated expectations help streamline data quality assurance processes, ensuring users can quickly and effectively apply industry-standard checks to their data assets.


# Custom Expectations

**How to Create a Expectation Template**

Navigate to Template Tab to Create an Expectation Template

<figure><img src="/files/kBneRfVyxQEEa7lt2JbB" alt=""><figcaption><p>Expectation Template</p></figcaption></figure>

Click on the + sign at the bottom right corner

Add Expectation statement to generate SQL Query

Select the Template Type as Expectation Template&#x20;

Check the details once the SQL statement is generated

Click on Submit to Continue

<figure><img src="/files/NZU01GtCmUFdveNRfmwN" alt=""><figcaption><p>Expectation Statement</p></figcaption></figure>

Templated Expectation has been created


# Profiling

VDA provides feature of data profiling , this involves reviewing data to better understand its structure and maintain data quality standards within an organization. The main purpose of this feature is to gain insights into the quality of the data by using methods to review, summarize, and evaluate its condition.

Data engineers typically perform this work using a range of business rules and analytical algorithms. VDA's data profiling evaluates data based on factors such as accuracy, consistency, and timeliness, identifying issues like inconsistencies, inaccuracies, or null values.

#### How to profile Data using VDA

Navigate to Data Source tab and click on the desired Data Source to find the Data Sets

<figure><img src="/files/AcVVFAEgR39LX84bvO3j" alt=""><figcaption><p>Data Sources Tab</p></figcaption></figure>

From the list of Data Sets click on the required Data Set

<figure><img src="/files/OeZ7YqnE1bz5AFjhNFGE" alt=""><figcaption><p>Data sets Page</p></figcaption></figure>

<figure><img src="/files/D0Xkf9EpNGJyPyxkkmc9" alt=""><figcaption><p>Details of a Dataset</p></figcaption></figure>

Select the checkbox next to the column name to run the Data Profiling

<figure><img src="/files/ke5xwvzH2JykfvsHN7hl" alt=""><figcaption><p>Column Selection to Run Data Profiling</p></figcaption></figure>

Click on the Graph icon positioned at bottom right corner of the screen to run the Profiling

Once the profiling is done the column name will be highlighted in blue Click on it and you will find different parameters based on the column data type or go to Profiling tab to check profiling details for a dataset .

<figure><img src="/files/Kwk7pWvidXaZvfqmFzxg" alt=""><figcaption><p>Profiling Completed</p></figcaption></figure>

#### Some Examples

Text/ String Data Type

#### Distinct Value

* **Definition**: The number of unique values present in a text field.
* **Importance**: Provides insights into data variety and uniqueness. High distinct values indicate diverse textual content, while low distinct values may suggest repetitive or standardized data entries.

#### Missing Value

* **Definition**: The count or percentage of records where the text field is empty or null.
* **Importance**: Indicates data completeness. High missing values can affect analysis and decision-making, highlighting potential gaps in data collection or entry processes.

#### Distinct Values

* **Definition**: The number of unique values present in the string field.
* **Importance**: Provides insights into data variety and uniqueness. High distinct values suggest diverse textual content, while low distinct values may indicate repetitive or standardized data entries.

#### Missing Values

* **Definition**: The count or percentage of records where the string field is empty or null.
* **Importance**: Indicates data completeness. High occurrences of missing values can affect analysis and decision-making, highlighting potential gaps in data collection or entry processes.

#### Total Characters

* **Definition**: The total number of characters (including spaces and special characters) across all string values in the field.
* **Importance**: Helps in understanding data volume and storage requirements. Larger total character counts may impact system performance and storage costs.

#### Unique Values Count

* **Definition**: The count of values that appear only once within the string field.
* **Importance**: Identifies truly unique entries, which can be critical for data deduplication and ensuring data accuracy.

#### Duplicate Values

* **Definition**: The count of values that appear more than once within the string field.
* **Importance**: Indicates data redundancy and potential data quality issues. Identifying and managing duplicates is essential for maintaining data integrity.

#### Lower Case Letters

* **Definition**: The count or percentage of characters in the string field that are in lower case.
* **Importance**: Provides insights into text normalization and consistency. Monitoring lower case usage helps in standardizing data for analysis and reporting purposes.

#### Upper Case Letters

* **Definition**: The count or percentage of characters in the string field that are in upper case.
* **Importance**: Similar to lower case letters, tracking upper case usage assists in data standardization and consistency checks.

#### Punctuations

* **Definition**: The count or percentage of punctuation characters (e.g., periods, commas, exclamation marks) within the string field.
* **Importance**: Helps in analyzing text complexity and identifying patterns in punctuation usage that may influence data processing and analysis.

<figure><img src="/files/lv1WwHbFdfWj4CaegCmZ" alt=""><figcaption><p>Data Profiling Example</p></figcaption></figure>

<figure><img src="/files/kN48a849m3uwT9MQHtMu" alt=""><figcaption><p>Data Profiling Example</p></figcaption></figure>

If Datatype is Number

#### Distinct Values

* **Definition**: The number of unique numeric values present in the field.
* **Importance**: Provides insights into data diversity and granularity. High distinct values indicate a wide range of data points, while low distinct values may suggest categorical or heavily aggregated data.

#### Missing Values

* **Definition**: The count or percentage of records where the numeric field is empty or null.
* **Importance**: Indicates data completeness. High occurrences of missing values can impact analysis and decision-making, necessitating data cleansing or imputation.

#### Negative Values

* **Definition**: The count or percentage of numeric values that are negative.
* **Importance**: Helps in understanding data trends and distributions. Negative values may be critical in specific contexts such as financial or scientific datasets.

#### Quantile Statistics

* **Definition**: Values that divide a dataset into equal portions, providing insights into data distribution.
* **Importance**: Helps in understanding data spread and variability. Common quantiles include quartiles (dividing data into quarters) and percentiles (dividing data into hundredths).

#### Descriptive Statistics

* **Definition**: Statistical summaries such as mean, median, mode, standard deviation, and variance.
* **Importance**: Provides a comprehensive view of central tendency, dispersion, and shape of the numeric data distribution.

#### Common Values and Frequencies

* **Definition**: The most frequently occurring numeric values and their occurrence count.
* **Importance**: Identifies popular or dominant values within the dataset, highlighting potential data trends or biases.

#### Maximum and Minimum Values

* **Definition**: The highest and lowest numeric values observed in the dataset.
* **Importance**: Indicates data range and extremes. Examining maximum and minimum values helps in identifying outliers or unusual data points that may require further investigation.

<figure><img src="/files/cdFJJf0VgPzKPjxH4rxE" alt=""><figcaption><p>Data Profiling for Number Datatype</p></figcaption></figure>

<figure><img src="/files/Sqp56IEOlXur5GXUNk9e" alt=""><figcaption><p>Data Profiling for Number Datatype</p></figcaption></figure>

<figure><img src="/files/uth62oMhvPl04Hxhtfz4" alt=""><figcaption><p>Data Profiling for Number Datatype</p></figcaption></figure>


# Reconciliation


# Data Analytics


# Data Modeling

Go to Datasource

<figure><img src="/files/HDatFgQ7oTTidzyQBR3M" alt=""><figcaption></figcaption></figure>

Click on the Dataset

<figure><img src="/files/M6QwaywMusoucbZvwgbz" alt=""><figcaption></figcaption></figure>

Generate the Metadata (icon on the bottom right corner)

<figure><img src="/files/JoEmUKon3vBdDKprWkrF" alt=""><figcaption></figcaption></figure>

Add short description of all the columns&#x20;

<figure><img src="/files/odPL8KgDDZpeYvWnxY9U" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/YwVVz7Llbs0V1Uuu0Q0u" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/7xKpTNNHePfUirZDUD0V" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/qlp57fkTmNuOPYHiIUKO" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/THWAUuT6H3Y3FcyOFVNL" alt=""><figcaption></figcaption></figure>

Go to Transformation

<figure><img src="/files/NLBsFErTRxWBQ0hsPI3C" alt=""><figcaption></figcaption></figure>

Add the transformation logic

<figure><img src="/files/wteus5ivqCuJfuj5b4Jq" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/l5H53zja59baDDlFRxaw" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/dKjvc72epJagvG7xwXLM" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/PqRAEEy5Y62ovjcRm5LV" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/tfdtPgYnzsCelp25CnSt" alt=""><figcaption></figcaption></figure>


# Visualization

Data analytics involves examining datasets to draw conclusions about the information they contain. This process utilizes various tools, techniques, and methodologies to identify patterns, trends, and relationships within the data. By transforming raw data into meaningful insights, VDA data analytics supports data-driven decision-making across diverse industries.

Data visualization is a crucial aspect of data analytics, as it translates complex data sets into graphical representations, making it easier to comprehend and interpret. VDA enable analysts to quickly identify trends, patterns, and outliers that may not be apparent from raw data alone. By presenting data visually, stakeholders can grasp the significance of their data insights more efficiently, facilitating better decision-making.

## Steps to create Visualization

Create the Datasource

<figure><img src="/files/EUbEJOxlrVIFkQyFfvvV" alt=""><figcaption><p>Datasource creation</p></figcaption></figure>

Click on Dashboard

Click on + icon (bottom right corner) to add a new Dashboard

<figure><img src="/files/e0zDrSwb6mk0E0yfn7am" alt=""><figcaption><p>Create Dashboard</p></figcaption></figure>

### Create Dashboard

Provide a Title

Select the data source

Write a description

<figure><img src="/files/Zoyf8MIYW2l5RNNhtdag" alt=""><figcaption><p>Create Dashboard</p></figcaption></figure>

<figure><img src="/files/BT1ylthcKLopbRrE5Xnx" alt=""><figcaption><p>Create Dashboard</p></figcaption></figure>

Click on the Dashboard which we have created

<figure><img src="/files/gr6pVDDFg1NNqiwrcPss" alt=""><figcaption><p>Create Dashboard</p></figcaption></figure>

Click on + icon to add KPI

Write a query in English and submit, SQL will be generated automatically

Example: Top 10 customer from Customer table

<figure><img src="/files/0rj9owFxNxE0WuOQea0T" alt=""><figcaption><p>Add KPI</p></figcaption></figure>

VDA will generate a SQL Query and return relevant result in the side panel

<figure><img src="/files/YPOU4T6YxTmlAdTH4hcF" alt=""><figcaption><p>Add KPI</p></figcaption></figure>

&#x20;Add a Title and write the required type of Chart

<figure><img src="/files/wIhsdDSNmrP84nHtpg7H" alt=""><figcaption><p>Add Chart Type</p></figcaption></figure>

<figure><img src="/files/vWsUuCRAxqq5PWUV8nBI" alt=""><figcaption><p>Visualization</p></figcaption></figure>

VDA supports a wide range of visualizations to explore and present data effectively. Here are the types of graphs and visualizations that can be created using VDA:

1. Bar Chart: Displays categorical data with rectangular bars, where the length of each bar corresponds to the value it represents.
2. Line Chart: Shows data points connected by straight line segments, commonly used to display trends over time.
3. Area Chart: Similar to a line chart but with the area below the line filled in, making it easy to visualize cumulative totals over time.
4. Pie Chart: Displays data as slices of a circular pie, where each slice represents a category's proportion of the whole.
5. Scatter Plot: Plots data points on a two-dimensional plane, useful for visualizing relationships between two variables.
6. Bubble Chart: A variation of a scatter plot where data points are represented as bubbles with varying sizes, useful for displaying three dimensions of data (x-axis, y-axis, and size).
7. Histogram: Shows the distribution of numerical data by dividing the data into bins and displaying bars of frequency counts.
8. Heatmap: Visualizes data using colors to represent values in a matrix, ideal for identifying patterns or correlations in large datasets.
9. Box Plot: Summarizes the distribution of numerical data through quartiles, providing insights into variability and outliers.
10. Time Series Chart: Specifically designed for time-based data, displaying trends and patterns over a period.
11. Sunburst Chart: A hierarchical chart that shows relationships between data categories through nested rings, useful for displaying hierarchical data structures.
12. Treemap: Visualizes hierarchical data using nested rectangles, where the size of each rectangle represents a quantitative measure.
13. Dual Axis Chart: Combines two different chart types with shared axes, allowing for easy comparison between two sets of data.
14. Sankey Diagram: Illustrates flows and relationships between different entities, showing the magnitude of flows through the width of paths.
15. Chord Diagram: Displays relationships between entities, visualizing the flow or connections between them through arcs.

&#x20;


# Data Ingestion

**Data Ingestion** is the process of collecting, importing, and processing data from various sources into a centralized data repository or system, making it ready for analysis and utilization. This step is critical in any data pipeline as it ensures that data is available and accessible in the desired format for subsequent processing, analysis, and decision-making.

#### Key Components of Data Ingestion

1. **Source Systems**
   * **Definition**: The origins of the data being ingested. These can include databases, APIs, flat files, cloud storage, sensors, and more.
   * **Variety**: Data can come from structured sources like SQL databases, semi-structured sources like JSON files, or unstructured sources like text documents.

VDA connectors:

**List of Available Connectors**

* [Amazon Athena](https://aws.amazon.com/athena/)
* [Amazon EventBridge](https://aws.amazon.com/eventbridge/)
* [Amazon Glue](https://aws.amazon.com/glue/) and anything built over it
* [Amazon Redshift](https://aws.amazon.com/redshift/)
* [Apache Cassandra](https://cassandra.apache.org/)
* [Apache Druid](https://druid.apache.org/)
* [Apache Hive](https://hive.apache.org/)
* CSV
* [dbt](https://www.getdbt.com/)
* [Delta Lake](https://delta.io/)
* [Elasticsearch](https://www.elastic.co/)
* [Google BigQuery](https://cloud.google.com/bigquery)
* [IBM DB2](https://www.ibm.com/analytics/db2)
* [Kafka Schema Registry](https://docs.confluent.io/platform/current/schema-registry/index.html)
* [Microsoft SQL Server](https://www.microsoft.com/en-us/sql-server/default.aspx)
* [MySQL](https://www.mysql.com/)
* [Oracle](https://www.oracle.com/index.html) (through dbapi or sql\_alchemy)
* [PostgreSQL](https://www.postgresql.org/)
* [PrestoDB](http://prestodb.io/)
* [Trino (formerly Presto SQL)](https://trino.io/)
* [Vertica](https://www.vertica.com/)
* [Snowflake](https://www.snowflake.com/)

Create the Data Source from which data is to be Ingested

<figure><img src="/files/QJjSxN4Yh8F5yXYvAFtt" alt=""><figcaption><p>Data Source Creation</p></figcaption></figure>

Navigate to Datasource Tab and click on the desired Data source to find the list of associated Datasets

<figure><img src="/files/tBzpZHz7gb2x0R4ULEf1" alt=""><figcaption><p>Datasets</p></figcaption></figure>

For New Ingestion Workbook creation, Navigate to Workbook, click on Ingestion Book create&#x20;

<figure><img src="/files/jx7BaEZezO2S8PsGXBM1" alt=""><figcaption><p>Create New Ingestion Book</p></figcaption></figure>

1. **Ingestion Methods**
   * **Batch Processing**: Data is collected and processed in large chunks at scheduled intervals.

     * **Use Cases**: Suitable for use cases where real-time data is not necessary, such as end-of-day reports or periodic data archiving.&#x20;
     * It can be:
       * Full refresh
       * Incremental
       * Historical

     <figure><img src="/files/sd5qiWyBUmOeBAf8aZOe" alt=""><figcaption><p>Batch Ingestion</p></figcaption></figure>

To create a Schedule, Navigate to Schedule, click on Plus and enter the details of Ingestion Workbook and Submit:

Name of Schedule

Name of Ingestion workbook

Frequency

Start Date

<figure><img src="/files/vnR0IL51ljxTUbGIatL0" alt=""><figcaption><p>Create a schedule</p></figcaption></figure>


# Governance

Data governance is the practice of managing and controlling data to ensure its accuracy, quality, security, and accessibility throughout its lifecycle. Effective data governance establishes policies, procedures, and standards to manage data assets, ensuring they are reliable and trusted. It plays a critical role in compliance, risk management, and decision-making processes within an organization.

**The Importance of Role-Based Access Control (RBAC) in Data Governance**

Role-Based Access Control (RBAC) is a method of regulating access to data based on the roles of individual users within an organization. By assigning permissions to specific roles rather than individual users, RBAC simplifies the management of user privileges and enhances data security. Integrating RBAC into data governance ensures that only authorized personnel can access, modify, or manage data according to their roles and responsibilities.

### **Implementing RBAC in VDA**

Login as SUPERADMIN user

Move to Permissions Tab

In the Permissions Page, Select the "Role Name" then from the below dropdown API list select the "Mappings Applies Selected Role" and click on submit.

<figure><img src="/files/QJipK5RF3YLK5z39i8PN" alt=""><figcaption><p>Access Management</p></figcaption></figure>

Select the appropriate permissions and then Click on Submit&#x20;

<figure><img src="/files/uyiMCldpbp1gg7leTkv3" alt=""><figcaption><p>Access Management</p></figcaption></figure>

After this, users should only be able to see pages for which they are authorized.

Ex. Datasource

<figure><img src="/files/ZelyYKmgwNU9Ost1FxlK" alt=""><figcaption><p>View</p></figcaption></figure>

### **Key Benefits**

1. **Enhanced Security**: By restricting data access based on roles, RBAC minimizes the risk of unauthorized access and potential data breaches.
2. **Compliance and Auditability**: RBAC helps organizations comply with regulatory requirements by providing clear audit trails of who accessed and modified data.
3. **Improved Efficiency**: Simplifies the process of managing user permissions, especially in large organizations with complex data environments.
4. **Scalability**: Facilitates easier administration of access controls as organizations grow and roles evolve.
5. **Risk Mitigation**: Reduces the likelihood of human error and misuse of data by ensuring that users only have access to data necessary for their roles.


