Python SDK
Official Python SDK for interacting with the PromptLayer API.
Installation
SDK evals
Define evals withevaluate(...), then run them with promptlayer eval run <path>. Start from the Quickstart or the SDK Evals overview.
Table scorecards
Useclient.tables.sheets.scorecards to configure criteria, run calculations, and inspect row-level results.
get(table_id, sheet_id)configure(table_id, sheet_id, body)delete(table_id, sheet_id)migrate_legacy_score(table_id, sheet_id, body)recalculate(table_id, sheet_id, body=None)cancel(table_id, sheet_id, body=None)get_calculation(table_id, sheet_id, calculation_id)list_rows(table_id, sheet_id, options=None)get_row(table_id, sheet_id, row_index, options=None)
PATCH to configure scorecards, returns 202 Accepted when recalculation is queued, paginates rows with cursor / next_cursor, and returns row breakdown fields at the top level.
Using the run Method (Recommended)
The easiest way to use PromptLayer is with the run() method. It fetches a prompt template from the Prompt Registry, executes it against your configured LLM provider, and logs the result — all in one call.
Your LLM API keys (OpenAI, Anthropic, etc.) are never sent to our servers. All LLM requests are made locally from your machine, PromptLayer just logs the request.
run() method works with any provider configured in your prompt template — OpenAI, Anthropic, Google, and more. See the Run documentation for full details.
After making your first few requests, you should be able to see them in the PromptLayer dashboard!

Basic Usage
For any LLM provider you plan to use, you must set its corresponding API key as an environment variable (for example,
OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_API_KEY etc.). The PromptLayer client does not support passing these keys directly in code. If the relevant environment variables are not set, any requests to those LLM providers will fail.Provider-Specific Configuration
Provider-Specific Configuration
Using Gemini models through Vertex AI
Python SDK: Set these environment variables:GOOGLE_GENAI_USE_VERTEXAI=trueGOOGLE_CLOUD_PROJECT="<google_cloud_project_id>"GOOGLE_CLOUD_LOCATION="region"GOOGLE_APPLICATION_CREDENTIALS="path/to/google_service_account_file.json"
Using Claude models through Vertex AI
Python SDK: Set these environment variables:ANTHROPIC_VERTEX_PROJECT_ID="<google_cloud_project_id>"CLOUD_ML_REGION="region"GOOGLE_APPLICATION_CREDENTIALS="path/to/google_service_account_file.json"
Python
Parameters
prompt_name/promptName(str, required): The name of the prompt to run.prompt_version/promptVersion(int, optional): Specific version of the prompt to use.prompt_release_label/promptReleaseLabel(str, optional): Release label of the prompt (e.g., “prod”, “staging”).input_variables/inputVariables(Dict[str, Any], optional): Variables to be inserted into the prompt template.tags(List[str], optional): Tags to associate with this run.metadata(Dict[str, str], optional): Additional metadata for the run.model_parameter_overrides/modelParameterOverrides(Union[Dict[str, Any], None], optional): Model-specific parameter overrides.stream(bool, default=False): Whether to stream the response.provider(str, optional): The LLM provider to use (e.g., “openai”, “anthropic”, “google”). This is useful if you want to override the provider specified in the prompt template.model(str, optional): The model to use (e.g., “gpt-4o”, “claude-3-7-sonnet-latest”, “gemini-2.5-flash”). This is useful if you want to override the model specified in the prompt template.
Return Value
The method returns a dictionary (Python) or object (JavaScript) with the following keys:request_id: Unique identifier for the request.raw_response: The raw response from the LLM provider.prompt_blueprint: The prompt blueprint used for the request.
Advanced Usage
Streaming
To stream the response:Python
prompt_blueprint, allowing you to track how the response is constructed in real-time. The request_id is only included in the final chunk.
Using Different Versions or Release Labels
Python
Adding Tags and Metadata
Python
Overriding Model Parameters
You can also overrideprovider and model at runtime to choose a different LLM provider or model. This is useful if you want to use a different provider than the one specified in the prompt template. PromptLayer will automatically return the correct llm_kwargs for the specified provider and model with default values for the parameters corresponding to the provider and model.
Python SDK
Running Workflows
Userun_workflow() to execute a PromptLayer Workflow from the Python SDK. Workflows are multi-step pipelines that can combine prompt, tool, code, and conditional nodes.
Python
Workflow Parameters
workflow_id_or_name(str or int, required): The Workflow name or ID to run.input_variables(Dict[str, Any], optional): Variables to pass into the Workflow.metadata(Dict[str, str], optional): Metadata to attach to the Workflow run.workflow_label_name(str, optional): Label name for the Workflow version, such as"production".workflow_version(int, optional): Specific Workflow version number to run.return_all_outputs(bool, default=False): Whether to return outputs for every Workflow node.timeout(int or float, optional): Maximum time, in seconds, to wait for the Workflow to complete.
workflow_name is still supported for backward compatibility, but workflow_id_or_name is the preferred parameter.Workflow Return Value
By default,run_workflow() returns the final output node’s value. When return_all_outputs=True, it returns a dictionary keyed by node name, including each node’s status, value, errors, and whether the node is an output node.
Python
return_all_outputs=True:
AsyncPromptLayer. See Async Workflow Execution for an async example.
SDK Cache
The PromptLayer Python SDK supports an in-memory template cache to reduce fetch latency and improve resilience when the PromptLayer API has transient failures. Enable cache when you want to:- Reduce repeated template fetch latency
- Lower dependency on real-time PromptLayer API availability
- Continue serving recently known-good templates during temporary API issues
cache_ttl_seconds when creating a client:
How It Works
When cache is enabled,templates.get() and run() use this flow:
- Return a fresh cached template if available.
- If cache is stale or missing, fetch from API and refresh cache.
- If API fetch fails with a transient error and a stale template exists, serve the stale template.
Stale fallback only applies to transient API errors (for example, timeout, connection, or internal server errors).
Important Behavior
- Cache is in-memory and process-local (not shared across machines/containers).
- Requests with
metadata_filtersormodel_parameter_overridesbypass cache. - Publishing via
templates.publish()invalidates cache for that prompt name.
Practical Guidance
- Start with
cache_ttl_secondsbetween60and300. - Use a shorter TTL if your prompts change frequently.
- Use a longer TTL if your prompts are stable and lower latency matters most.
- Keep
throw_on_error=Trueif you want hard failures when no cache entry is available.
Custom Logging with log_request
If you need more control — for example, using your own LLM client, a custom provider, or background processing — you can use log_request to manually log requests to PromptLayer.
Error Handling
PromptLayer provides robust error handling with specialized exception classes and configurable error behavior.Exception Classes
The library includes specific exception types following industry best practices:Using throw_on_error
By default, PromptLayer throws exceptions when errors occur. You can control this behavior using the throw_on_error parameter:
Automatic Retry Mechanism
PromptLayer includes a built-in retry mechanism to handle transient failures gracefully. This ensures your application remains resilient when temporary issues occur. Retry Behavior:- Total Attempts: 4 attempts (1 initial + 3 retries)
- Exponential Backoff: Retries wait progressively longer between attempts (2s, 4s, 8s)
- Max Wait Time: 15 seconds maximum wait between retries
- 5xx Server Errors: Internal server errors, service unavailable, etc.
- 429 Rate Limit Errors: When rate limits are exceeded
- Connection Errors: Network connectivity issues
- Timeout Errors: Request timeouts
- 4xx Client Errors (except 429): Bad requests, authentication errors, not found, etc.
The retry mechanism operates transparently in the background. You don’t need to implement retry logic yourself - PromptLayer handles it automatically for recoverable errors.
Logging
PromptLayer uses Python’s built-inlogging module for all log output:
Async Support
PromptLayer supports asynchronous operations, ideal for managing concurrent tasks in non-blocking environments like web servers, microservices, or Jupyter notebooks.Initializing the Async Client
To use asynchronous non-blocking methods, initialize AsyncPromptLayer as shown:Async Usage Examples
The asynchronous client functions similarly to the synchronous version, but allows for non-blocking execution withasyncio. Below are example uses.
Example 1: Async Template Management
Use asynchronous methods to manage templates:Example 2: Async Workflow Execution
Run Workflows asynchronously for better efficiency:Example 3: Async Tracking and Logging
Track and log requests asynchronously:Example 4: Asynchronous Prompt Execution with run Method
You can execute prompt templates asynchronously using the run method. This allows you to run a prompt template by name with given input variables.Example 5: Asynchronous Streaming Prompt Execution with run Method
You can run streaming prompt template using the run method as well.prompt_blueprint, allowing you to track how the response is constructed in real-time.
Want to say hi 👋, submit a feature request, or report a bug? ✉️ Contact us

