> ## Documentation Index
> Fetch the complete documentation index at: https://docs.pretectum.io/llms.txt
> Use this file to discover all available pages before exploring further.

# List Datasets

> Retrieve the list of datasets within a specific schema

The List Datasets endpoint returns all datasets defined within a specific schema. Datasets are collections of data objects that share the same schema structure and represent actual data records in your master data repository.

## Prerequisites

* A Pretectum API key (see [API Keys](/api-reference/authentication/api-keys))
* Permission to access datasets in your tenant
* Valid business area ID (see [List Business Areas](/api-reference/business-areas/list))
* Valid schema ID (see [List Schemas](/api-reference/schemas/list))

## Authentication

Include your API key in the `Authorization` header.

```bash theme={null}
Authorization: pre_your_api_key
```

## Request

### Path Parameters

<ParamField path="businessAreaId" type="string" required>
  The unique identifier of the business area. You can obtain this from the [List Business Areas](/api-reference/business-areas/list) endpoint.
</ParamField>

<ParamField path="schemaId" type="string" required>
  The unique identifier of the schema containing the datasets. You can obtain this from the [List Schemas](/api-reference/schemas/list) endpoint.
</ParamField>

### Query Parameters

<ParamField query="pageKey" type="string">
  A pagination token for retrieving the next page of results. This value is returned in the response as `nextPageKey` when more results are available.
</ParamField>

### Headers

<ParamField header="Authorization" type="string" required initialValue="pre_your_api_key">
  Your Pretectum API key. Create one in the Pretectum app under **Configuration → API Keys**.
</ParamField>

<ParamField header="Accept" type="string" default="application/json">
  The response content type. Currently only `application/json` is supported.
</ParamField>

### Example Requests

<CodeGroup>
  ```bash cURL theme={null}
  # List all datasets in a schema
  curl -X GET "https://api.pretectum.io/v1/businessareas/20240115103000123a1b2c3d4e5f6789012345678901234/schemas/20240115103000456d1e2f3a4b5c6789012345678901234/datasets" \
    -H "Authorization: pre_your_api_key" \
    -H "Accept: application/json"

  # Paginate through datasets
  curl -X GET "https://api.pretectum.io/v1/businessareas/20240115103000123a1b2c3d4e5f6789012345678901234/schemas/20240115103000456d1e2f3a4b5c6789012345678901234/datasets?pageKey=eyJMYXN0RXZhbHVhdGVkS2V5Ijp7Li4ufQ" \
    -H "Authorization: pre_your_api_key" \
    -H "Accept: application/json"
  ```

  ```javascript JavaScript theme={null}
  const apiKey = 'pre_your_api_key';
  const businessAreaId = '20240115103000123a1b2c3d4e5f6789012345678901234';
  const schemaId = '20240115103000456d1e2f3a4b5c6789012345678901234';

  async function getDatasets(businessAreaId, schemaId, pageKey = null) {
    const url = new URL(
      `https://api.pretectum.io/v1/businessareas/${businessAreaId}/schemas/${schemaId}/datasets`
    );
    if (pageKey) {
      url.searchParams.set('pageKey', pageKey);
    }

    const response = await fetch(url, {
      headers: {
        'Authorization': apiKey,
        'Accept': 'application/json'
      }
    });

    return response.json();
  }

  const datasets = await getDatasets(businessAreaId, schemaId);
  console.log(`Found ${datasets.items.length} datasets`);
  datasets.items.forEach(dataset => {
    console.log(`- ${dataset.dataSetName}: ${dataset.recordCount} records`);
  });
  ```

  ```python Python theme={null}
  import requests

  api_key = 'pre_your_api_key'
  business_area_id = '20240115103000123a1b2c3d4e5f6789012345678901234'
  schema_id = '20240115103000456d1e2f3a4b5c6789012345678901234'

  def get_datasets(business_area_id, schema_id, page_key=None):
      params = {}
      if page_key:
          params['pageKey'] = page_key

      response = requests.get(
          f'https://api.pretectum.io/v1/businessareas/{business_area_id}/schemas/{schema_id}/datasets',
          params=params,
          headers={
              'Authorization': api_key,
              'Accept': 'application/json'
          }
      )
      response.raise_for_status()
      return response.json()

  datasets = get_datasets(business_area_id, schema_id)
  print(f"Found {len(datasets['items'])} datasets")
  for dataset in datasets['items']:
      print(f"- {dataset['dataSetName']}: {dataset['recordCount']} records")
  ```
</CodeGroup>

## Response

A successful request returns an object containing an array of datasets and pagination information.

<ResponseField name="items" type="array" required>
  An array of dataset objects within the schema.

  <Expandable title="Dataset object properties">
    <ResponseField name="dataSetId" type="string">
      The unique identifier for the dataset. Use this ID when filtering data object searches.
    </ResponseField>

    <ResponseField name="dataSetName" type="string">
      The display name of the dataset. This is the human-readable name you can use in the `dataSet` filter parameter when searching data objects.
    </ResponseField>

    <ResponseField name="dataSetDescription" type="string">
      A description of the dataset explaining its purpose and the data it contains.
    </ResponseField>

    <ResponseField name="businessAreaId" type="string">
      The ID of the business area this dataset belongs to.
    </ResponseField>

    <ResponseField name="businessAreaName" type="string">
      The name of the business area this dataset belongs to.
    </ResponseField>

    <ResponseField name="schemaId" type="string">
      The ID of the schema this dataset uses.
    </ResponseField>

    <ResponseField name="schemaName" type="string">
      The name of the schema this dataset uses.
    </ResponseField>

    <ResponseField name="recordCount" type="integer">
      The total number of data objects (records) in this dataset.
    </ResponseField>

    <ResponseField name="erroredRecordsCount" type="integer">
      The number of records that have validation errors.
    </ResponseField>

    <ResponseField name="runningJobsCount" type="integer">
      The number of background jobs currently running on this dataset (e.g., imports, exports).
    </ResponseField>

    <ResponseField name="version" type="integer">
      The version number of the dataset. This increments each time the dataset is modified.
    </ResponseField>

    <ResponseField name="createdBy" type="string">
      The identifier of the user who created the dataset.
    </ResponseField>

    <ResponseField name="createdByEmail" type="string">
      The email address of the user who created the dataset.
    </ResponseField>

    <ResponseField name="createdByName" type="string">
      The full name of the user who created the dataset.
    </ResponseField>

    <ResponseField name="updatedBy" type="string">
      The identifier of the user who last modified the dataset.
    </ResponseField>

    <ResponseField name="updatedByEmail" type="string">
      The email address of the user who last modified the dataset.
    </ResponseField>

    <ResponseField name="updatedByName" type="string">
      The full name of the user who last modified the dataset.
    </ResponseField>

    <ResponseField name="createdDate" type="string">
      The ISO 8601 timestamp when the dataset was created.
    </ResponseField>

    <ResponseField name="updatedDate" type="string">
      The ISO 8601 timestamp when the dataset was last modified.
    </ResponseField>

    <ResponseField name="deleted" type="boolean">
      Indicates whether the dataset has been marked as deleted.
    </ResponseField>
  </Expandable>
</ResponseField>

<ResponseField name="nextPageKey" type="string">
  A pagination token for retrieving the next page of results. If this field is present, more datasets are available. Pass this value as the `pageKey` query parameter in your next request.
</ResponseField>

### Example Response

```json theme={null}
{
  "items": [
    {
      "dataSetId": "20240925152201042a1b2c3d4e5f6789012345678901234",
      "dataSetName": "US Customers",
      "dataSetDescription": "Customer records for United States region",
      "businessAreaId": "20240115103000123a1b2c3d4e5f6789012345678901234",
      "businessAreaName": "Customer",
      "schemaId": "20240115103000456d1e2f3a4b5c6789012345678901234",
      "schemaName": "Individual Customer",
      "recordCount": 15420,
      "erroredRecordsCount": 12,
      "runningJobsCount": 0,
      "version": 5,
      "createdBy": "9ae5f422-bb62-4c9d-b277-594ddcda6d8d",
      "createdByEmail": "admin@example.com",
      "createdByName": "John Admin",
      "updatedBy": "9ae5f422-bb62-4c9d-b277-594ddcda6d8d",
      "updatedByEmail": "admin@example.com",
      "updatedByName": "John Admin",
      "createdDate": "2024-09-25T15:22:01.042Z",
      "updatedDate": "2024-12-15T10:30:00.000Z",
      "deleted": false
    },
    {
      "dataSetId": "20240926090000123b2c3d4e5f6a7890123456789012345",
      "dataSetName": "European Customers",
      "dataSetDescription": "Customer records for European region",
      "businessAreaId": "20240115103000123a1b2c3d4e5f6789012345678901234",
      "businessAreaName": "Customer",
      "schemaId": "20240115103000456d1e2f3a4b5c6789012345678901234",
      "schemaName": "Individual Customer",
      "recordCount": 8750,
      "erroredRecordsCount": 3,
      "runningJobsCount": 0,
      "version": 2,
      "createdBy": "b5f6g733-cc73-5d0e-c388-605eeda7e9e",
      "createdByEmail": "data_manager@example.com",
      "createdByName": "Jane Manager",
      "updatedBy": "b5f6g733-cc73-5d0e-c388-605eeda7e9e",
      "updatedByEmail": "data_manager@example.com",
      "updatedByName": "Jane Manager",
      "createdDate": "2024-09-26T09:00:00.123Z",
      "updatedDate": "2024-11-20T14:45:00.000Z",
      "deleted": false
    }
  ],
  "nextPageKey": "eyJMYXN0RXZhbHVhdGVkS2V5Ijp7ImRhdGFTZXRJZCI6IjIwMjQwOTI2MDkwMDAw..."
}
```

### Response Without Pagination

When all datasets fit in a single response, no `nextPageKey` is returned:

```json theme={null}
{
  "items": [
    {
      "dataSetId": "20240925152201042a1b2c3d4e5f6789012345678901234",
      "dataSetName": "US Customers",
      "dataSetDescription": "Customer records for United States region",
      "businessAreaId": "20240115103000123a1b2c3d4e5f6789012345678901234",
      "businessAreaName": "Customer",
      "schemaId": "20240115103000456d1e2f3a4b5c6789012345678901234",
      "schemaName": "Individual Customer",
      "recordCount": 15420,
      "erroredRecordsCount": 0,
      "runningJobsCount": 0,
      "version": 5,
      "createdBy": "9ae5f422-bb62-4c9d-b277-594ddcda6d8d",
      "createdByEmail": "admin@example.com",
      "createdByName": "John Admin",
      "updatedBy": "9ae5f422-bb62-4c9d-b277-594ddcda6d8d",
      "updatedByEmail": "admin@example.com",
      "updatedByName": "John Admin",
      "createdDate": "2024-09-25T15:22:01.042Z",
      "updatedDate": "2024-12-15T10:30:00.000Z",
      "deleted": false
    }
  ]
}
```

### Empty Response

If the schema has no datasets defined, the response will contain an empty items array:

```json theme={null}
{
  "items": []
}
```

## Error Responses

| Status Code | Description |
| - | - |
| `401 Unauthorized` | The API key is missing, malformed, unknown, inactive, expired or deleted. Check the key in **Configuration → API Keys**. |
| `403 Forbidden` | Your application client does not have permission to access datasets. Contact your tenant administrator. |
| `404 Not Found` | The specified business area or schema does not exist, or you do not have access to it. |
| `500 Internal Server Error` | An unexpected error occurred on the server. Try again later or contact support. |

## Pagination

When a schema contains many datasets, results are paginated. Use the `nextPageKey` from the response to fetch subsequent pages:

<CodeGroup>
  ```javascript JavaScript theme={null}
  async function getAllDatasets(businessAreaId, schemaId) {
    const allDatasets = [];
    let pageKey = null;

    do {
      const response = await getDatasets(businessAreaId, schemaId, pageKey);
      allDatasets.push(...response.items);
      pageKey = response.nextPageKey;
    } while (pageKey);

    return allDatasets;
  }

  const allDatasets = await getAllDatasets(businessAreaId, schemaId);
  console.log(`Total datasets: ${allDatasets.length}`);
  ```

  ```python Python theme={null}
  def get_all_datasets(business_area_id, schema_id):
      all_datasets = []
      page_key = None

      while True:
          response = get_datasets(business_area_id, schema_id, page_key)
          all_datasets.extend(response['items'])

          page_key = response.get('nextPageKey')
          if not page_key:
              break

      return all_datasets

  all_datasets = get_all_datasets(business_area_id, schema_id)
  print(f"Total datasets: {len(all_datasets)}")
  ```
</CodeGroup>

## Use Cases

### Filtering Search Results by Dataset

Use the dataset names returned by this endpoint to filter your data object searches:

```bash theme={null}
# First, get the list of datasets
curl -X GET "https://api.pretectum.io/v1/businessareas/{businessAreaId}/schemas/{schemaId}/datasets" \
  -H "Authorization: pre_your_api_key"

# Then search within a specific dataset
curl -X GET "https://api.pretectum.io/v1/dataobjects/search?query=John&dataSet=US%20Customers" \
  -H "Authorization: pre_your_api_key"
```

### Monitoring Data Quality

Check the `erroredRecordsCount` to identify datasets with data quality issues:

```javascript theme={null}
const datasets = await getDatasets(businessAreaId, schemaId);

const datasetsWithErrors = datasets.items.filter(ds => ds.erroredRecordsCount > 0);
datasetsWithErrors.forEach(ds => {
  console.log(`${ds.dataSetName}: ${ds.erroredRecordsCount} errors out of ${ds.recordCount} records`);
});
```

### Tracking Record Counts

Monitor the size of your datasets:

```python theme={null}
datasets = get_datasets(business_area_id, schema_id)

total_records = sum(ds['recordCount'] for ds in datasets['items'])
print(f"Total records across all datasets: {total_records}")

for ds in sorted(datasets['items'], key=lambda x: x['recordCount'], reverse=True):
    print(f"  {ds['dataSetName']}: {ds['recordCount']:,} records")
```

## Best Practices

1. **Cache dataset lists**: Dataset metadata changes less frequently than record data. Cache the response and refresh periodically.
2. **Filter by active datasets**: Exclude datasets where `deleted: true` in user-facing interfaces.
3. **Use names for search filters**: When filtering searches with the `dataSet` parameter, use the `dataSetName` field value.
4. **Handle pagination**: Always check for `nextPageKey` in responses and fetch all pages if needed.
5. **Monitor error counts**: Regularly check `erroredRecordsCount` to identify data quality issues early.

## Related Endpoints

<CardGroup cols={2}>
  <Card title="List Schemas" icon="file-code" href="/api-reference/schemas/list">
    Get schema IDs for dataset queries
  </Card>

  <Card title="List Business Areas" icon="building" href="/api-reference/business-areas/list">
    Get business area IDs
  </Card>

  <Card title="Search Data Objects" icon="magnifying-glass" href="/api-reference/dataobjects/search">
    Search within specific datasets
  </Card>

  <Card title="API Keys" icon="key" href="/api-reference/authentication/api-keys">
    Obtain authentication token
  </Card>
</CardGroup>
