Last modified: Oct 10, 2026
How to Install Apache Airflow in Python
Apache Airflow is a workflow orchestration platform. It lets you schedule, monitor, and manage data pipelines with pure Python code. Data engineers use it to run ETL jobs, train models, and move files between systems.
This guide walks you through installing Airflow from scratch. You will learn the prerequisites, the exact pip command, how to initialize the database, and how to start the web server.
By the end, you will have a working Airflow instance and a simple DAG running on your machine.
What Is Apache Airflow?
Airflow is an open source platform built by Airbnb and now maintained by the Apache Software Foundation. It defines workflows as code using Python.
A workflow is called a DAG, which stands for Directed Acyclic Graph. Each node in the DAG is a task. Edges define the order in which tasks run.
The scheduler triggers tasks based on time or dependencies. The web server provides a UI for monitoring and manual runs. The metadata database stores DAG state, task history, and connections.
Airflow is not a data processing engine. It orchestrates other tools. You use it to call Python functions, run SQL queries, trigger Spark jobs, or send notifications.
Prerequisites Before You Install
Airflow has stricter requirements than most Python libraries. Check your environment first.
You need Python 3.8 to 3.11. Airflow does not yet support 3.12 in all versions. Verify your version.
python --version Python 3.11.6 You also need pip and a virtual environment tool. The venv module is built into Python, so no extra install is needed.
Important: Never install Airflow into your global Python. It has many dependencies that can conflict with system packages.
Create a Virtual Environment
Isolate your Airflow installation in a dedicated virtual environment. This is the recommended approach in the official docs.
python -m venv airflow-env source airflow-env/bin/activate On Windows, the activation command is different.
airflow-env\Scripts\activate Once activated, your prompt shows the environment name. Every package you install now stays inside that folder.
Set the Airflow Home Directory
Airflow stores its configuration, DAGs, and logs in a home directory. The default is ~/airflow. You can override it with an environment variable.
export AIRFLOW_HOME=~/airflow On Windows, use the set command instead.
set AIRFLOW_HOME=%USERPROFILE%\airflow Add this line to your shell profile so it persists across sessions. Airflow will create the folder automatically on first run.
Install Airflow with pip
The official installation uses a constraints file. The constraints lock dependency versions that are known to work together. Skipping them causes dependency conflicts.
Install with the constraints file for your Python version. Here is the command for Python 3.11.
AIRFLOW_VERSION=2.9.1 PYTHON_VERSION="$(python --version | cut -d " " -f 2 | cut -d "." -f 1-2)" CONSTRAINT_URL="https://raw.githubusercontent.com/apache/airflow/constraints-${AIRFLOW_VERSION}/constraints-${PYTHON_VERSION}.txt" pip install "apache-airflow==${AIRFLOW_VERSION}" --constraint "${CONSTRAINT_URL}" Pip downloads Airflow and its dependencies. This can take several minutes. The output looks like this.
Collecting apache-airflow==2.9.1 Downloading apache_airflow-2.9.1-py3-none-any.whl (6.2 MB) Collecting apache-airflow-core==2.9.1 Downloading apache_airflow_core-2.9.1-py3-none-any.whl (4.1 MB) ... Successfully installed apache-airflow-2.9.1 apache-airflow-core-2.9.1 ... Always pin the version. Airflow changes fast between releases. Pinning keeps your setup reproducible.
If you want to install a specific set of extras, add them to the package name.
pip install "apache-airflow[postgres,google]==2.9.1" --constraint "${CONSTRAINT_URL}" This installs the Postgres and Google provider packages alongside the core.
Initialize the Database
Airflow uses a metadata database. By default it uses SQLite, which is fine for local development.
Run the database initialization command. It creates the SQLite file and all required tables.
airflow db init DB: sqlite:////home/user/airflow/airflow.db [2024-05-01 10:00:00,000] {db.py:1234} INFO - Creating tables [2024-05-01 10:00:01,000] {db.py:1234} INFO - Initialization done Warning: SQLite is not supported for production. Use PostgreSQL or MySQL for real workloads. SQLite does not handle concurrent writes well.
Create an Admin User
You need a user account to log into the web UI. Create one with the airflow users create command.
airflow users create \ --username admin \ --firstname Admin \ --lastname User \ --role Admin \ --email admin@example.com Airflow prompts for a password. Choose a strong one. The role Admin grants full access to the UI and configuration.
Password: Repeat for confirmation: User "admin" created with role "Admin" Start the Web Server
The web server provides the Airflow UI. Start it in the foreground for a quick test.
airflow webserver --port 8080 ____________ _____________ ____ |__( )_________ __/__ /________ __ ____ /| |_ /__ ___/_ /_ __ /_ __ \_ | /| / / ___ ___ | / _ / _ __/ _ / / /_/ /_ |/ |/ / _/_/ |_/_/ /_/ /_/ /_/ /_/\____/_/|_/|__/ [2024-05-01 10:05:00,000] {webserver_command.py:123} INFO - Starting the web server on port 8080 Open your browser and navigate to http://localhost:8080. Log in with the admin account you created.
You will see the Airflow dashboard with an empty DAG list.
Start the Scheduler
The scheduler is the heart of Airflow. It parses DAGs and triggers tasks. Open a new terminal, activate the environment, and start it.
airflow scheduler [2024-05-01 10:06:00,000] {scheduler_job_runner.py:123} INFO - Starting the scheduler [2024-05-01 10:06:00,100] {scheduler_job_runner.py:456} INFO - Processing 0 files The scheduler runs continuously. Keep it open in a separate terminal while you develop DAGs.
Tip: Use a process manager like systemd or supervisor for production. Do not rely on manual terminal sessions.
Write Your First DAG
Create a file in the ~/airflow/dags folder. Airflow scans this folder for DAG definitions.
Here is a simple DAG with two Python tasks.
from datetime import datetime, timedelta from airflow import DAG from airflow.operators.python import PythonOperator # Default arguments applied to every task default_args = { 'owner': 'admin', 'retries': 1, 'retry_delay': timedelta(minutes=5), } # Define the DAG with DAG( dag_id='hello_airflow', default_args=default_args, description='A simple tutorial DAG', schedule_interval='@daily', start_date=datetime(2024, 5, 1), catchup=False, tags=['example'], ) as dag: def say_hello(): print("Hello from Airflow!") def say_goodbye(): print("Goodbye from Airflow!") # Define tasks task_hello = PythonOperator( task_id='say_hello', python_callable=say_hello, ) task_goodbye = PythonOperator( task_id='say_goodbye', python_callable=say_goodbye, ) # Set the order: hello runs first, then goodbye task_hello >> task_goodbye Save the file as hello_dag.py in the dags folder. The scheduler picks it up within a minute.
Refresh the web UI. The new DAG appears in the list. Toggle it on and trigger a manual run.
[2024-05-01 10:10:00,000] {python.py:177} INFO - Done. Returned value was: None [2024-05-01 10:10:00,100] {taskinstance.py:1234} INFO - Marking task as SUCCESS You can check the logs for each task in the UI. Click the task, then click "Logs" to see the output.
Common Errors and Troubleshooting
ModuleNotFoundError: No module named 'airflow' means you are not in the virtual environment. Activate it and try again.
Dependency conflict during install means you skipped the constraints file. Always include the --constraint flag with the correct URL.
Port 8080 already in use means another service is running there. Start the webserver with a different port.
airflow webserver --port 8081 DAG does not appear in the UI means the scheduler did not parse it. Check that the file is in the dags folder and that Python syntax is valid.
Airflow db init fails usually means the AIRFLOW_HOME folder is not writable. Check permissions on the directory.
Scheduler crashes on startup often means a bad DAG file. Run airflow dags list to see the parsing error.
Best Practices for Airflow Installations
Pin the Airflow version and always use the constraints file. This prevents surprise breakages.
Use PostgreSQL for the metadata database, even in development. It matches production and avoids SQLite quirks.
Set the executor based on your scale. SequentialExecutor works locally. LocalExecutor is fine for a single machine. CeleryExecutor handles clusters.
Store secrets in environment variables or a secrets backend. Never hardcode credentials in DAG files.
Keep DAG files small and idempotent. Complex logic belongs in separate modules that DAGs import.
Run the scheduler and webserver as managed services. Use systemd, supervisor, or a container orchestrator.
If your DAGs send alerts or reports by email, you can integrate a mail provider. For example, you can send HTML emails with SendGrid and Python from a task to notify your team when a pipeline finishes.
If your pipelines also generate reports, you can produce the PDF inside a task and email it as an attachment. A common pattern is to generate PDFs with ReportLab, then attach the result to a notification email.
Conclusion
Installing Apache Airflow in Python takes a few careful steps. Create a virtual environment. Set the AIRFLOW_HOME directory. Install with the official constraints file. Initialize the database and create an admin user.
Then start the web server and scheduler. Write a simple DAG and watch it run. That is the entire flow.
Airflow scales from a local development machine to a distributed production cluster. Follow the examples in this guide to get started, then explore providers, sensors, and custom operators as your pipelines grow.
Start with a small DAG today. Add retries, alerts, and integrations as you learn. Happy orchestrating!