azure-storage-file-datalake-py
Azure Data Lake Storage Gen2 SDK for Python. Use for hierarchical file systems, big data analytics, and file/directory operations. Triggers: "data lake", "DataLakeServiceClient", "FileSystemClient", "ADLS Gen2", "hierarchical namespace".
Install
npx skills add https://github.com/microsoft/skills/tree/main/.github/plugins/azure-sdk-python/skills/azure-storage-file-datalake-py
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install microsoft-skills@llmmart
git clone https://github.com/microsoft/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole microsoft/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Azure Data Lake Storage Gen2 SDK for Python
Hierarchical file system for big data analytics workloads.
Installation
pip install azure-storage-file-datalake azure-identity
Environment Variables
AZURE_STORAGE_ACCOUNT_URL=https://<account>.dfs.core.windows.net # Required for all auth methods
AZURE_TOKEN_CREDENTIALS=prod # Required only if DefaultAzureCredential is used in production
Authentication & Lifecycle
🔑 Two rules apply to every code sample below:
- Prefer
DefaultAzureCredential. It works locally (Azure CLI / VS Code / Developer CLI) and in Azure (managed identity, workload identity) with no code change. Avoid connection strings, account/API keys — they bypass Entra audit and rotation.
- Local dev:
DefaultAzureCredentialworks as-is.- Production: set
AZURE_TOKEN_CREDENTIALS=prod(orAZURE_TOKEN_CREDENTIALS=<specific_credential>) to constrain the credential chain to production-safe credentials.- Wrap every client in a context manager so HTTP transports, sockets, and token caches are released deterministically:
- Sync:
with <Client>(...) as client:- Async:
async with <Client>(...) as client:andasync with DefaultAzureCredential() as credential:(fromazure.identity.aio)Snippets may abbreviate this setup, but production code should always follow both rules.
from azure.identity import DefaultAzureCredential, ManagedIdentityCredential
from azure.storage.filedatalake import DataLakeServiceClient
# Local dev: DefaultAzureCredential. Production: set AZURE_TOKEN_CREDENTIALS=prod or AZURE_TOKEN_CREDENTIALS=<specific_credential>
credential = DefaultAzureCredential(require_envvar=True)
# Or use a specific credential directly in production:
# See https://learn.microsoft.com/python/api/overview/azure/identity-readme?view=azure-python#credential-classes
# credential = ManagedIdentityCredential()
account_url = "https://<account>.dfs.core.windows.net"
with DataLakeServiceClient(account_url=account_url, credential=credential) as service_client:
# Use service_client here (see following sections for operations)
...
Client Hierarchy
| Client | Purpose |
|---|---|
DataLakeServiceClient |
Account-level operations |
FileSystemClient |
Container (file system) operations |
DataLakeDirectoryClient |
Directory operations |
DataLakeFileClient |
File operations |
File System Operations
# Create file system (container)
file_system_client = service_client.create_file_system("myfilesystem")
# Get existing
file_system_client = service_client.get_file_system_client("myfilesystem")
# Delete
service_client.delete_file_system("myfilesystem")
# List file systems
for fs in service_client.list_file_systems():
print(fs.name)
Directory Operations
file_system_client = service_client.get_file_system_client("myfilesystem")
# Create directory
directory_client = file_system_client.create_directory("mydir")
# Create nested directories
directory_client = file_system_client.create_directory("path/to/nested/dir")
# Get directory client
directory_client = file_system_client.get_directory_client("mydir")
# Delete directory
directory_client.delete_directory()
# Rename/move directory
directory_client.rename_directory(new_name="myfilesystem/newname")
File Operations
Upload File
# Get file client
file_client = file_system_client.get_file_client("path/to/file.txt")
# Upload from local file
with open("local-file.txt", "rb") as data:
file_client.upload_data(data, overwrite=True)
# Upload bytes
file_client.upload_data(b"Hello, Data Lake!", overwrite=True)
# Append data (for large files)
file_client.append_data(data=b"chunk1", offset=0, length=6)
file_client.append_data(data=b"chunk2", offset=6, length=6)
file_client.flush_data(12) # Commit the data
Download File
file_client = file_system_client.get_file_client("path/to/file.txt")
# Download all content
download = file_client.download_file()
content = download.readall()
# Download to file
with open("downloaded.txt", "wb") as f:
download = file_client.download_file()
download.readinto(f)
# Download range
download = file_client.download_file(offset=0, length=100)
Delete File
file_client.delete_file()
List Contents
# List paths (files and directories)
for path in file_system_client.get_paths():
print(f"{'DIR' if path.is_directory else 'FILE'}: {path.name}")
# List paths in directory
for path in file_system_client.get_paths(path="mydir"):
print(path.name)
# Recursive listing
for path in file_system_client.get_paths(path="mydir", recursive=True):
print(path.name)
File/Directory Properties
# Get properties
properties = file_client.get_file_properties()
print(f"Size: {properties.size}")
print(f"Last modified: {properties.last_modified}")
# Set metadata
file_client.set_metadata(metadata={"processed": "true"})
Access Control (ACL)
# Get ACL
acl = directory_client.get_access_control()
print(f"Owner: {acl['owner']}")
print(f"Permissions: {acl['permissions']}")
# Set ACL
directory_client.set_access_control(
owner="user-id",
permissions="rwxr-x---"
)
# Update ACL entries
from azure.storage.filedatalake import AccessControlChangeResult
directory_client.update_access_control_recursive(
acl="user:user-id:rwx"
)
Async Client
from azure.storage.filedatalake.aio import DataLakeServiceClient
from azure.identity.aio import DefaultAzureCredential
async def datalake_operations():
async with DefaultAzureCredential() as credential:
async with DataLakeServiceClient(
account_url="https://<account>.dfs.core.windows.net",
credential=credential
) as service_client:
file_system_client = service_client.get_file_system_client("myfilesystem")
file_client = file_system_client.get_file_client("test.txt")
await file_client.upload_data(b"async content", overwrite=True)
download = await file_client.download_file()
content = await download.readall()
import asyncio
asyncio.run(datalake_operations())
Best Practices
- Pick sync OR async and stay consistent. Do not mix
azure.storage.filedatalakesync clients withazure.storage.filedatalake.aioasync clients in the same call path. Choose one mode per module. - Always use context managers for clients and async credentials. Wrap every client in
with DataLakeServiceClient(...) as client:(sync) orasync with DataLakeServiceClient(...) as client:(async). For asyncDefaultAzureCredentialfromazure.identity.aio, also useasync with credential:so tokens and transports are cleaned up. - Use
DefaultAzureCredentialfor portable auth across local dev and Azure (avoid connection strings / API keys when possible). - Use hierarchical namespace for file system semantics
- Use
append_data+flush_datafor large file uploads - Set ACLs at directory level and inherit to children
- Use async client for high-throughput scenarios
- Use
get_pathswithrecursive=Truefor full directory listing - Set metadata for custom file attributes
- Consider Blob API for simple object storage use cases
Reference Files
| File | Contents |
|---|---|
| references/capabilities.md | Additional non-hero capabilities, operation-group coverage, and production checklists. |
| references/non-hero-scenarios.md | Dedicated non-hero examples for secondary/advanced scenarios. |
Files (skills)
-
references
-
capabilities.md 1.6 KB
# azure-storage-file-datalake-py capability coverage **SDK/package**: `azure-storage-file-datalake` This index maps hero scenarios in `SKILL.md` and links non-hero scenarios documented in dedicated reference files. ## Hero scenarios covered in SKILL.md - `Client Hierarchy` - `File System Operations` - `Directory Operations` - `File Operations` ## Non-hero scenarios - `List Contents`: Dedicated example and implementation notes. See: [`non-hero-scenarios.md#list-contents`](non-hero-scenarios.md#list-contents) - `File/Directory Properties`: Dedicated example and implementation notes. See: [`non-hero-scenarios.md#filedirectory-properties`](non-hero-scenarios.md#filedirectory-properties) - `Access Control (ACL)`: Dedicated example and implementation notes. See: [`non-hero-scenarios.md#access-control-acl`](non-hero-scenarios.md#access-control-acl) - `Async Client`: Dedicated example and implementation notes. See: [`non-hero-scenarios.md#async-client`](non-hero-scenarios.md#async-client) ## Related deep-dive references - [`non-hero-scenarios.md`](non-hero-scenarios.md): Dedicated non-hero examples and implementation notes. ## API breadth checklist - Verify client/auth mode for the environment before coding. - Confirm operation-group/method names against current Microsoft Learn API reference. - For Python SDKs with both sync and async clients, document both forms without a blanket preference. - Include cleanup/delete paths for created resources in examples. - Prefer idempotent create/update operations where available. - Validate paging/LRO/error-handling patterns for production paths. -
non-hero-scenarios.md 2.1 KB
# azure-storage-file-datalake-py non-hero scenarios These scenarios are intentionally separate from hero flows in `SKILL.md`. They cover secondary/advanced patterns typically used after the primary end-to-end path is working. ## List Contents ```python # List paths (files and directories) for path in file_system_client.get_paths(): print(f"{'DIR' if path.is_directory else 'FILE'}: {path.name}") # List paths in directory for path in file_system_client.get_paths(path="mydir"): print(path.name) # Recursive listing for path in file_system_client.get_paths(path="mydir", recursive=True): print(path.name) ``` ## File/Directory Properties ```python # Get properties properties = file_client.get_file_properties() print(f"Size: {properties.size}") print(f"Last modified: {properties.last_modified}") # Set metadata file_client.set_metadata(metadata={"processed": "true"}) ``` ## Access Control (ACL) ```python # Get ACL acl = directory_client.get_access_control() print(f"Owner: {acl['owner']}") print(f"Permissions: {acl['permissions']}") # Set ACL directory_client.set_access_control( owner="user-id", permissions="rwxr-x---" ) # Update ACL entries from azure.storage.filedatalake import AccessControlChangeResult directory_client.update_access_control_recursive( acl="user:user-id:rwx" ) ``` ## Async Client ```python from azure.storage.filedatalake.aio import DataLakeServiceClient from azure.identity.aio import DefaultAzureCredential async def datalake_operations(): async with DefaultAzureCredential() as credential: async with DataLakeServiceClient( account_url="https://<account>.dfs.core.windows.net", credential=credential ) as service_client: file_system_client = service_client.get_file_system_client("myfilesystem") file_client = file_system_client.get_file_client("test.txt") await file_client.upload_data(b"async content", overwrite=True) download = await file_client.download_file() content = await download.readall() import asyncio asyncio.run(datalake_operations()) ```
-
-
SKILL.md 8 KB
--- name: azure-storage-file-datalake-py description: | Azure Data Lake Storage Gen2 SDK for Python. Use for hierarchical file systems, big data analytics, and file/directory operations. Triggers: "data lake", "DataLakeServiceClient", "FileSystemClient", "ADLS Gen2", "hierarchical namespace". license: MIT metadata: author: Microsoft version: "1.0.0" package: azure-storage-file-datalake --- # Azure Data Lake Storage Gen2 SDK for Python Hierarchical file system for big data analytics workloads. ## Installation ```bash pip install azure-storage-file-datalake azure-identity ``` ## Environment Variables ```bash AZURE_STORAGE_ACCOUNT_URL=https://<account>.dfs.core.windows.net # Required for all auth methods AZURE_TOKEN_CREDENTIALS=prod # Required only if DefaultAzureCredential is used in production ``` ## Authentication & Lifecycle > **🔑 Two rules apply to every code sample below:** > > 1. **Prefer `DefaultAzureCredential`.** It works locally (Azure CLI / VS Code / Developer CLI) and in Azure (managed identity, workload identity) with no code change. Avoid connection strings, account/API keys — they bypass Entra audit and rotation. > - Local dev: `DefaultAzureCredential` works as-is. > - Production: set `AZURE_TOKEN_CREDENTIALS=prod` (or `AZURE_TOKEN_CREDENTIALS=<specific_credential>`) to constrain the credential chain to production-safe credentials. > 2. **Wrap every client in a context manager** so HTTP transports, sockets, and token caches are released deterministically: > - Sync: `with <Client>(...) as client:` > - Async: `async with <Client>(...) as client:` **and** `async with DefaultAzureCredential() as credential:` (from `azure.identity.aio`) > > Snippets may abbreviate this setup, but production code should always follow both rules. ```python from azure.identity import DefaultAzureCredential, ManagedIdentityCredential from azure.storage.filedatalake import DataLakeServiceClient # Local dev: DefaultAzureCredential. Production: set AZURE_TOKEN_CREDENTIALS=prod or AZURE_TOKEN_CREDENTIALS=<specific_credential> credential = DefaultAzureCredential(require_envvar=True) # Or use a specific credential directly in production: # See https://learn.microsoft.com/python/api/overview/azure/identity-readme?view=azure-python#credential-classes # credential = ManagedIdentityCredential() account_url = "https://<account>.dfs.core.windows.net" with DataLakeServiceClient(account_url=account_url, credential=credential) as service_client: # Use service_client here (see following sections for operations) ... ``` ## Client Hierarchy | Client | Purpose | |--------|---------| | `DataLakeServiceClient` | Account-level operations | | `FileSystemClient` | Container (file system) operations | | `DataLakeDirectoryClient` | Directory operations | | `DataLakeFileClient` | File operations | ## File System Operations ```python # Create file system (container) file_system_client = service_client.create_file_system("myfilesystem") # Get existing file_system_client = service_client.get_file_system_client("myfilesystem") # Delete service_client.delete_file_system("myfilesystem") # List file systems for fs in service_client.list_file_systems(): print(fs.name) ``` ## Directory Operations ```python file_system_client = service_client.get_file_system_client("myfilesystem") # Create directory directory_client = file_system_client.create_directory("mydir") # Create nested directories directory_client = file_system_client.create_directory("path/to/nested/dir") # Get directory client directory_client = file_system_client.get_directory_client("mydir") # Delete directory directory_client.delete_directory() # Rename/move directory directory_client.rename_directory(new_name="myfilesystem/newname") ``` ## File Operations ### Upload File ```python # Get file client file_client = file_system_client.get_file_client("path/to/file.txt") # Upload from local file with open("local-file.txt", "rb") as data: file_client.upload_data(data, overwrite=True) # Upload bytes file_client.upload_data(b"Hello, Data Lake!", overwrite=True) # Append data (for large files) file_client.append_data(data=b"chunk1", offset=0, length=6) file_client.append_data(data=b"chunk2", offset=6, length=6) file_client.flush_data(12) # Commit the data ``` ### Download File ```python file_client = file_system_client.get_file_client("path/to/file.txt") # Download all content download = file_client.download_file() content = download.readall() # Download to file with open("downloaded.txt", "wb") as f: download = file_client.download_file() download.readinto(f) # Download range download = file_client.download_file(offset=0, length=100) ``` ### Delete File ```python file_client.delete_file() ``` ## List Contents ```python # List paths (files and directories) for path in file_system_client.get_paths(): print(f"{'DIR' if path.is_directory else 'FILE'}: {path.name}") # List paths in directory for path in file_system_client.get_paths(path="mydir"): print(path.name) # Recursive listing for path in file_system_client.get_paths(path="mydir", recursive=True): print(path.name) ``` ## File/Directory Properties ```python # Get properties properties = file_client.get_file_properties() print(f"Size: {properties.size}") print(f"Last modified: {properties.last_modified}") # Set metadata file_client.set_metadata(metadata={"processed": "true"}) ``` ## Access Control (ACL) ```python # Get ACL acl = directory_client.get_access_control() print(f"Owner: {acl['owner']}") print(f"Permissions: {acl['permissions']}") # Set ACL directory_client.set_access_control( owner="user-id", permissions="rwxr-x---" ) # Update ACL entries from azure.storage.filedatalake import AccessControlChangeResult directory_client.update_access_control_recursive( acl="user:user-id:rwx" ) ``` ## Async Client ```python from azure.storage.filedatalake.aio import DataLakeServiceClient from azure.identity.aio import DefaultAzureCredential async def datalake_operations(): async with DefaultAzureCredential() as credential: async with DataLakeServiceClient( account_url="https://<account>.dfs.core.windows.net", credential=credential ) as service_client: file_system_client = service_client.get_file_system_client("myfilesystem") file_client = file_system_client.get_file_client("test.txt") await file_client.upload_data(b"async content", overwrite=True) download = await file_client.download_file() content = await download.readall() import asyncio asyncio.run(datalake_operations()) ``` ## Best Practices 1. **Pick sync OR async and stay consistent.** Do not mix `azure.storage.filedatalake` sync clients with `azure.storage.filedatalake.aio` async clients in the same call path. Choose one mode per module. 2. **Always use context managers for clients and async credentials.** Wrap every client in `with DataLakeServiceClient(...) as client:` (sync) or `async with DataLakeServiceClient(...) as client:` (async). For async `DefaultAzureCredential` from `azure.identity.aio`, also use `async with credential:` so tokens and transports are cleaned up. 3. **Use `DefaultAzureCredential`** for portable auth across local dev and Azure (avoid connection strings / API keys when possible). 4. **Use hierarchical namespace** for file system semantics 5. **Use `append_data` + `flush_data`** for large file uploads 6. **Set ACLs at directory level** and inherit to children 7. **Use async client** for high-throughput scenarios 8. **Use `get_paths` with `recursive=True`** for full directory listing 9. **Set metadata** for custom file attributes 10. **Consider Blob API** for simple object storage use cases ## Reference Files | File | Contents | |------|----------| | [references/capabilities.md](references/capabilities.md) | Additional non-hero capabilities, operation-group coverage, and production checklists. | | [references/non-hero-scenarios.md](references/non-hero-scenarios.md) | Dedicated non-hero examples for secondary/advanced scenarios. |
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.