{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "d04d0d74-acc0-49b1-8ca2-0c88909dd73d",
   "metadata": {},
   "source": [
    "# OMI–Benzene Enrichment Pipeline\n",
    "\n",
    "This script enriches AMTIC benzene observations with collocated **OMI satellite columns** (NO₂, tropospheric NO₂, O₃, SO₂, HCHO) using pre-computed pixel indices for each monitoring site. It is designed for **speed and repeatability** on NCAR Derecho by avoiding on-the-fly nearest-neighbor searches and by directly indexing the HDF5 arrays.\n",
    "\n",
    "## What the pipeline does\n",
    "\n",
    "At a high level, the workflow loads a site→satellite **index map** (row/column positions on the OMI grid for each SITEID), scans your OMI L2G folders to find the right file for each measurement date, and then extracts the per-orbit pixel values directly at those indices. It aggregates valid orbital measurements into daily means, carries along simple quality/orbit counts, and writes the enriched table back to CSV.\n",
    "\n",
    "## Inputs and outputs\n",
    "\n",
    "The pipeline expects three key locations:\n",
    "\n",
    "- **Index file** (pre-computed):  \n",
    "  `/glade/derecho/scratch/shahriar/research_projects/benzene_prediction/01_raw_data/epa_amtic/benzene_sites_satellite_indices.csv`  \n",
    "  Each row maps a `SITEID` to OMI grid indices (`OMI_LAT_INDEX`, `OMI_LON_INDEX`) and an approximate distance in km.\n",
    "\n",
    "- **Benzene (base) file**:  \n",
    "  `/glade/derecho/scratch/shahriar/research_projects/benzene_prediction/01_raw_data/epa_amtic/benzene_hcho_amtic_regional.csv`  \n",
    "  Contains at least `SITEID`, `STATE`, `DATE`, and `BENZENE`. The script adds satellite columns to this table.\n",
    "\n",
    "- **OMI satellite base directory** (subfolders per species):  \n",
    "  `/glade/derecho/scratch/shahriar/research_projects/benzene_prediction/01_raw_data/satellite/`  \n",
    "  Required subpaths and file patterns (one file per day/species):\n",
    "  - `omi_no2/OMI-Aura_L2G-OMNO2G_YYYYmMMDD_vXXX-*.he5`\n",
    "  - `omi_o3/OMI-Aura_L2G-OMTO3G_YYYYmMMDD_vXXX-*.he5`\n",
    "  - `omi_so2/OMI-Aura_L2G-OMSO2G_YYYYmMMDD_vXXX-*.he5`\n",
    "  - `omi_hcho/OMI-Aura_L2G-OMHCHOG_YYYYmMMDD_vXXX-*.he5`\n",
    "\n",
    "The **output file** is:\n",
    "`/glade/derecho/scratch/shahriar/research_projects/benzene_prediction/01_raw_data/epa_amtic/benzene_with_omi_data_1.csv`\n",
    "\n",
    "It includes the original benzene rows plus columns:\n",
    "\n",
    "- `NO2`, `NO2_QUALITY`, `NO2_N_ORBITS`, `NO2_TROP`\n",
    "- `O3`, `O3_QUALITY`, `O3_N_ORBITS`\n",
    "- `SO2`, `SO2_QUALITY`, `SO2_N_ORBITS`\n",
    "- `HCHO`, `HCHO_QUALITY`, `HCHO_N_ORBITS`\n",
    "\n",
    "## How the code is structured\n",
    "\n",
    "The pipeline is organized into small, focused functions:\n",
    "\n",
    "- **`load_index_mapping()`**  \n",
    "  Reads the precomputed index file, prints how many sites were found, and builds a fast dictionary keyed by `SITEID` with `omi_lat_idx`, `omi_lon_idx`, and `omi_distance`. If the file is missing, it exits early with an informative message.\n",
    "\n",
    "- **`extract_date_from_filename(filename)`**  \n",
    "  A helper to parse a date from an OMI filename using a regex. This is illustrative and not central to the main control flow because the pipeline builds search patterns by date rather than parsing file names first.\n",
    "\n",
    "- **`find_omi_file(satellite_base, gas_type, target_date)`**  \n",
    "  Converts the `YYYY-MM-DD` date into the OMI pattern `YYYYmMMDD`, constructs a glob pattern for the correct **gas folder + daily file**, and returns the first matching path if it exists. If a daily file is missing, the script just skips that gas and moves on gracefully.\n",
    "\n",
    "- **`extract_omi_value_fast(filepath, gas_type, site_id, indices_dict)`**  \n",
    "  The performance-critical routine. It opens the `.he5` once with `h5py`, then **directly indexes** `main_data[:, lat_idx, lon_idx]` to get all orbital values for the day at that pixel. It filters out fill values (`-1.2676506e+30`) and NaNs, computes the mean over valid orbits, and counts `n_valid_orbits`. For NO₂, it also extracts the **tropospheric** column at the same indices. It attempts to read a quality flag array; if that dataset is unavailable for a species, it safely defaults.\n",
    "\n",
    "- **`extract_all_omi_data()`**  \n",
    "  Orchestrates the run. It loads benzene and indices, initializes empty columns for each gas (including quality/orbit counts), and converts `DATE` to a proper datetime. It iterates over **unique dates** (skipping pre-launch days before 2004-10-01), finds daily OMI files per gas, and then iterates over the benzene records for that date, populating the satellite fields in place. It reports progress (dates processed, rate, ETA), prints a compact summary at the end, and saves the enriched CSV.\n",
    "\n",
    "## Data extraction details\n",
    "\n",
    "For each `(SITEID, DATE)` pair:\n",
    "1. Look up the OMI **grid indices** for that site.\n",
    "2. For each species that has a file for that day, read the **3-D array** from the appropriate HDF5 path:\n",
    "   - NO₂: `HDFEOS/GRIDS/ColumnAmountNO2/Data Fields/ColumnAmountNO2`\n",
    "   - NO₂ (trop): `HDFEOS/GRIDS/ColumnAmountNO2/Data Fields/ColumnAmountNO2Trop`\n",
    "   - O₃: `HDFEOS/GRIDS/OMI Column Amount O3/Data Fields/ColumnAmountO3`\n",
    "   - SO₂: `HDFEOS/GRIDS/OMI Total Column Amount SO2/Data Fields/ColumnAmountSO2`\n",
    "   - HCHO: `HDFEOS/GRIDS/OMI Total Column Amount HCHO/Data Fields/ColumnAmountHCHO`\n",
    "3. Slice all orbits at the station pixel: `data[:, lat_idx, lon_idx]`.\n",
    "4. Remove fill values and NaNs, then take the **mean** over valid orbits.\n",
    "5. Save the mean to the species column, the number of valid orbits to `*_N_ORBITS`, the quality code to `*_QUALITY`, and for NO₂ also save `NO2_TROP`.\n",
    "\n",
    "This approach assumes the index file has already chosen the “best” pixel for each site (nearest-neighbor on the OMI grid). It deliberately avoids repeated geodesic nearest-pixel searches to maximize throughput.\n",
    "\n",
    "## Performance characteristics\n",
    "\n",
    "The main speedups come from:\n",
    "- **Precomputed pixel indices** (no per-record spatial search).\n",
    "- **Date-wise batching** (open each OMI file once per day).\n",
    "- **Direct HDF5 indexing** across the orbit dimension.\n",
    "- **Minimal I/O** (skip days without files; skip gases without files; only read needed datasets).\n",
    "\n",
    "You’ll see progress logs every ~50 dates with **records/sec** and a simple ETA. The final summary reports total records, species-wise successful extractions, and total runtime.\n",
    "\n",
    "## Assumptions and safeguards\n",
    "\n",
    "- `SITEID` in benzene and in the index file uses the same identifiers.\n",
    "- The OMI directory structure and filename patterns follow the standard L2G naming. If a date’s file is missing, that gas is skipped for that date without failing the run.\n",
    "- Fill value handling uses the documented sentinel `-1.2676506e+30`; quality arrays are read when available and ignored otherwise.\n",
    "- Dates before **2004-10-01** (pre-OMI) are automatically skipped.\n",
    "\n",
    "## How to run\n",
    "\n",
    "From a Derecho login node or interactive job with the correct modules/conda env:\n",
    "\n",
    "```bash\n",
    "python extract_omi_for_benzene.py\n",
    "```\n",
    "\n",
    "If you placed the function block in a notebook or another script, make sure the three paths in `extract_all_omi_data()` (satellite base, benzene file, output file) match your storage layout.\n",
    "\n",
    "## Verifying results\n",
    "\n",
    "After the run, spot-check a few rows by `(SITEID, DATE)`:\n",
    "- Confirm that days without OMI files remain `NaN`.\n",
    "- Verify that `*_N_ORBITS` is reasonable (typically up to the number of daily orbits).\n",
    "- Compare a small random sample against an independent tool (e.g., HDFView) by navigating to the same grid indices and confirming the mean of non-fill values.\n",
    "\n",
    "## Extending or adapting\n",
    "\n",
    "- To change the aggregation (e.g., median instead of mean), replace the `np.mean(valid_values)` line.\n",
    "- To use a **quality mask** rather than reporting raw flags, apply a boolean selection based on your project’s QA criteria before computing means.\n",
    "- To support additional gases with L2G structure, add their HDF5 paths to `data_paths` and an entry in `gas_configs`."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "acf428b3-e663-4a52-ac6e-ea090a519998",
   "metadata": {},
   "outputs": [],
   "source": [
    "import pandas as pd\n",
    "import numpy as np\n",
    "import h5py\n",
    "import os\n",
    "import glob\n",
    "from datetime import datetime\n",
    "import re\n",
    "import warnings\n",
    "\n",
    "warnings.filterwarnings('ignore')\n",
    "\n",
    "def load_index_mapping():\n",
    "    \"\"\"Load the pre-computed satellite indices\"\"\"\n",
    "    \n",
    "    indices_file = '/glade/derecho/scratch/shahriar/research_projects/benzene_prediction/01_raw_data/epa_amtic/benzene_sites_satellite_indices.csv'\n",
    "    \n",
    "    if not os.path.exists(indices_file):\n",
    "        print(f\"ERROR: Index file not found: {indices_file}\")\n",
    "        print(\"You need to run the index generation first!\")\n",
    "        return None\n",
    "    \n",
    "    indices_df = pd.read_csv(indices_file)\n",
    "    print(f\"Loaded indices for {len(indices_df)} sites\")\n",
    "    \n",
    "    # Convert to dictionary for fast lookup\n",
    "    indices_dict = {}\n",
    "    for _, row in indices_df.iterrows():\n",
    "        indices_dict[row['SITEID']] = {\n",
    "            'omi_lat_idx': row['OMI_LAT_INDEX'],\n",
    "            'omi_lon_idx': row['OMI_LON_INDEX'],\n",
    "            'omi_distance': row['OMI_DISTANCE_KM']\n",
    "        }\n",
    "    \n",
    "    return indices_dict\n",
    "\n",
    "def extract_date_from_filename(filename):\n",
    "    \"\"\"Extract date from OMI filename\"\"\"\n",
    "    pattern = r'_(\\d{4})m(\\d{2})(\\d{2})_v\\d{3}'\n",
    "    match = re.search(pattern, filename)\n",
    "    if match:\n",
    "        year, month, day = match.groups()\n",
    "        return f\"{year}-{month.zfill(2)}-{day.zfill(2)}\"\n",
    "    return None\n",
    "\n",
    "def find_omi_file(satellite_base, gas_type, target_date):\n",
    "    \"\"\"Find OMI file for specific gas and date\"\"\"\n",
    "    \n",
    "    gas_configs = {\n",
    "        'NO2': 'omi_no2/OMI-Aura_L2G-OMNO2G_*m*.he5',\n",
    "        'O3': 'omi_o3/OMI-Aura_L2G-OMTO3G_*m*.he5', \n",
    "        'SO2': 'omi_so2/OMI-Aura_L2G-OMSO2G_*m*.he5',\n",
    "        'HCHO': 'omi_hcho/OMI-Aura_L2G-OMHCHOG_*m*.he5'\n",
    "    }\n",
    "    \n",
    "    # Convert date to filename format\n",
    "    date_obj = datetime.strptime(target_date, '%Y-%m-%d')\n",
    "    date_pattern = f\"{date_obj.year}m{date_obj.month:02d}{date_obj.day:02d}\"\n",
    "    \n",
    "    # Search for file\n",
    "    pattern = gas_configs[gas_type].replace('*m*', f'*{date_pattern}*')\n",
    "    search_path = os.path.join(satellite_base, pattern)\n",
    "    files = glob.glob(search_path)\n",
    "    \n",
    "    return files[0] if files else None\n",
    "\n",
    "def extract_omi_value_fast(filepath, gas_type, site_id, indices_dict):\n",
    "    \"\"\"Ultra-fast extraction using pre-computed indices\"\"\"\n",
    "    \n",
    "    data_paths = {\n",
    "        'NO2': {\n",
    "            'main': 'HDFEOS/GRIDS/ColumnAmountNO2/Data Fields/ColumnAmountNO2',\n",
    "            'trop': 'HDFEOS/GRIDS/ColumnAmountNO2/Data Fields/ColumnAmountNO2Trop',\n",
    "            'quality': 'HDFEOS/GRIDS/ColumnAmountNO2/Data Fields/VcdQualityFlags'\n",
    "        },\n",
    "        'O3': {\n",
    "            'main': 'HDFEOS/GRIDS/OMI Column Amount O3/Data Fields/ColumnAmountO3',\n",
    "            'quality': 'HDFEOS/GRIDS/OMI Column Amount O3/Data Fields/QualityFlags'\n",
    "        },\n",
    "        'SO2': {\n",
    "            'main': 'HDFEOS/GRIDS/OMI Total Column Amount SO2/Data Fields/ColumnAmountSO2',\n",
    "            'quality': 'HDFEOS/GRIDS/OMI Total Column Amount SO2/Data Fields/GroundPixelQualityFlags'\n",
    "        },\n",
    "        'HCHO': {\n",
    "            'main': 'HDFEOS/GRIDS/OMI Total Column Amount HCHO/Data Fields/ColumnAmountHCHO',\n",
    "            'quality': 'HDFEOS/GRIDS/OMI Total Column Amount HCHO/Data Fields/MainDataQualityFlag'\n",
    "        }\n",
    "    }\n",
    "    \n",
    "    # Get pre-computed indices\n",
    "    if site_id not in indices_dict:\n",
    "        return {'value': np.nan, 'trop_value': np.nan, 'quality': np.nan, 'n_orbits': 0}\n",
    "    \n",
    "    lat_idx = indices_dict[site_id]['omi_lat_idx']\n",
    "    lon_idx = indices_dict[site_id]['omi_lon_idx']\n",
    "    \n",
    "    # Skip if no valid indices\n",
    "    if lat_idx < 0 or lon_idx < 0:\n",
    "        return {'value': np.nan, 'trop_value': np.nan, 'quality': np.nan, 'n_orbits': 0}\n",
    "    \n",
    "    try:\n",
    "        with h5py.File(filepath, 'r') as f:\n",
    "            paths = data_paths[gas_type]\n",
    "            \n",
    "            # Extract main data - direct index access!\n",
    "            main_data = f[paths['main']][...]\n",
    "            orbital_values = main_data[:, lat_idx, lon_idx]  # All 15 orbits\n",
    "            \n",
    "            # Filter valid values\n",
    "            fill_value = -1.2676506e+30\n",
    "            valid_values = orbital_values[orbital_values != fill_value]\n",
    "            valid_values = valid_values[~np.isnan(valid_values)]\n",
    "            \n",
    "            # Calculate mean of valid orbital measurements\n",
    "            main_value = np.mean(valid_values) if len(valid_values) > 0 else np.nan\n",
    "            n_valid_orbits = len(valid_values)\n",
    "            \n",
    "            # Extract tropospheric column for NO2\n",
    "            trop_value = np.nan\n",
    "            if gas_type == 'NO2' and 'trop' in paths:\n",
    "                try:\n",
    "                    trop_data = f[paths['trop']][...]\n",
    "                    trop_orbital = trop_data[:, lat_idx, lon_idx]\n",
    "                    valid_trop = trop_orbital[trop_orbital != fill_value]\n",
    "                    valid_trop = valid_trop[~np.isnan(valid_trop)]\n",
    "                    trop_value = np.mean(valid_trop) if len(valid_trop) > 0 else np.nan\n",
    "                except:\n",
    "                    pass\n",
    "            \n",
    "            # Extract quality flag\n",
    "            quality_flag = 0\n",
    "            try:\n",
    "                quality_data = f[paths['quality']][...]\n",
    "                quality_orbital = quality_data[:, lat_idx, lon_idx]\n",
    "                valid_quality = quality_orbital[quality_orbital != 65535]\n",
    "                quality_flag = int(valid_quality[0]) if len(valid_quality) > 0 else 0\n",
    "            except:\n",
    "                pass\n",
    "            \n",
    "            return {\n",
    "                'value': float(main_value) if not np.isnan(main_value) else np.nan,\n",
    "                'trop_value': float(trop_value) if not np.isnan(trop_value) else np.nan,\n",
    "                'quality': quality_flag,\n",
    "                'n_orbits': n_valid_orbits\n",
    "            }\n",
    "            \n",
    "    except Exception as e:\n",
    "        print(f\"  Error extracting {gas_type}: {e}\")\n",
    "        return {'value': np.nan, 'trop_value': np.nan, 'quality': np.nan, 'n_orbits': 0}\n",
    "\n",
    "def extract_all_omi_data():\n",
    "    \"\"\"Main function to extract OMI data for all benzene measurements\"\"\"\n",
    "    \n",
    "    print(\"OMI DATA EXTRACTION FOR BENZENE SITES\")\n",
    "    print(\"=\" * 60)\n",
    "    \n",
    "    # File paths\n",
    "    satellite_base = '/glade/derecho/scratch/shahriar/research_projects/benzene_prediction/01_raw_data/satellite/'\n",
    "    benzene_file = '/glade/derecho/scratch/shahriar/research_projects/benzene_prediction/01_raw_data/epa_amtic/benzene_hcho_amtic_regional.csv'\n",
    "    output_file = '/glade/derecho/scratch/shahriar/research_projects/benzene_prediction/01_raw_data/epa_amtic/benzene_with_omi_data_1.csv'\n",
    "    \n",
    "    # Load data\n",
    "    print(\"Loading data...\")\n",
    "    indices_dict = load_index_mapping()\n",
    "    if indices_dict is None:\n",
    "        return None\n",
    "    \n",
    "    df_benzene = pd.read_csv(benzene_file)\n",
    "    print(f\"Loaded {len(df_benzene):,} benzene measurements\")\n",
    "    \n",
    "    # Initialize new columns\n",
    "    gases = ['NO2', 'O3', 'SO2', 'HCHO']\n",
    "    for gas in gases:\n",
    "        df_benzene[gas] = np.nan\n",
    "        df_benzene[f'{gas}_QUALITY'] = np.nan\n",
    "        df_benzene[f'{gas}_N_ORBITS'] = np.nan\n",
    "    \n",
    "    # Add NO2 tropospheric column\n",
    "    df_benzene['NO2_TROP'] = np.nan\n",
    "    \n",
    "    # Process by unique dates for efficiency\n",
    "    df_benzene['DATE_PARSED'] = pd.to_datetime(df_benzene['DATE'])\n",
    "    unique_dates = df_benzene['DATE_PARSED'].dt.strftime('%Y-%m-%d').unique()\n",
    "    \n",
    "    print(f\"Processing {len(unique_dates)} unique dates...\")\n",
    "    print(\"Using ultra-fast index-based extraction!\")\n",
    "    \n",
    "    processed_records = 0\n",
    "    successful_extractions = {gas: 0 for gas in gases}\n",
    "    \n",
    "    import time\n",
    "    start_time = time.time()\n",
    "    \n",
    "    for i, date_str in enumerate(sorted(unique_dates)):\n",
    "        # Skip dates before OMI launch\n",
    "        if date_str < '2004-10-01':\n",
    "            continue\n",
    "        \n",
    "        # Progress reporting\n",
    "        if (i + 1) % 50 == 0 or i == 0:\n",
    "            elapsed = time.time() - start_time\n",
    "            rate = processed_records / elapsed if elapsed > 0 else 0\n",
    "            remaining_dates = len(unique_dates) - i - 1\n",
    "            eta = remaining_dates / (i + 1) * elapsed if i > 0 else 0\n",
    "            \n",
    "            print(f\"  Progress: {i+1:,}/{len(unique_dates):,} dates ({(i+1)/len(unique_dates)*100:.1f}%)\")\n",
    "            print(f\"  Processing rate: {rate:.1f} records/sec | ETA: {eta/60:.1f} min\")\n",
    "        \n",
    "        # Get records for this date\n",
    "        date_mask = df_benzene['DATE_PARSED'].dt.strftime('%Y-%m-%d') == date_str\n",
    "        date_records = df_benzene[date_mask]\n",
    "        \n",
    "        # Find satellite files for this date\n",
    "        satellite_files = {}\n",
    "        for gas in gases:\n",
    "            sat_file = find_omi_file(satellite_base, gas, date_str)\n",
    "            if sat_file and os.path.exists(sat_file):\n",
    "                satellite_files[gas] = sat_file\n",
    "        \n",
    "        if not satellite_files:\n",
    "            continue\n",
    "        \n",
    "        # Extract data for each site on this date\n",
    "        for idx, row in date_records.iterrows():\n",
    "            site_id = row['SITEID']\n",
    "            \n",
    "            # Extract for each available gas\n",
    "            for gas, filepath in satellite_files.items():\n",
    "                result = extract_omi_value_fast(filepath, gas, site_id, indices_dict)\n",
    "                \n",
    "                # Store results\n",
    "                df_benzene.loc[idx, gas] = result['value']\n",
    "                df_benzene.loc[idx, f'{gas}_QUALITY'] = result['quality']\n",
    "                df_benzene.loc[idx, f'{gas}_N_ORBITS'] = result['n_orbits']\n",
    "                \n",
    "                # Store tropospheric NO2\n",
    "                if gas == 'NO2':\n",
    "                    df_benzene.loc[idx, 'NO2_TROP'] = result['trop_value']\n",
    "                \n",
    "                if not np.isnan(result['value']):\n",
    "                    successful_extractions[gas] += 1\n",
    "            \n",
    "            processed_records += 1\n",
    "    \n",
    "    # Final timing\n",
    "    total_time = time.time() - start_time\n",
    "    \n",
    "    print(f\"\\nEXTRACTION COMPLETE!\")\n",
    "    print(f\"Total time: {total_time/60:.1f} minutes ({total_time/3600:.2f} hours)\")\n",
    "    print(f\"Processing rate: {processed_records/total_time:.1f} records/second\")\n",
    "    \n",
    "    # Results summary\n",
    "    print(f\"\\nEXTRACTION SUMMARY:\")\n",
    "    print(f\"Total records processed: {processed_records:,}\")\n",
    "    for gas in gases:\n",
    "        success_rate = successful_extractions[gas] / processed_records * 100 if processed_records > 0 else 0\n",
    "        print(f\"{gas}: {successful_extractions[gas]:,} successful ({success_rate:.1f}%)\")\n",
    "    \n",
    "    # Clean up\n",
    "    df_benzene = df_benzene.drop('DATE_PARSED', axis=1)\n",
    "    \n",
    "    # Save results\n",
    "    print(f\"\\nSaving results...\")\n",
    "    df_benzene.to_csv(output_file, index=False)\n",
    "    print(f\"Results saved to: {output_file}\")\n",
    "    \n",
    "    # Show sample\n",
    "    print(f\"\\nSample results:\")\n",
    "    sample_cols = ['SITEID', 'STATE', 'DATE', 'BENZENE'] + gases\n",
    "    available_cols = [col for col in sample_cols if col in df_benzene.columns]\n",
    "    print(df_benzene[available_cols].head().to_string())\n",
    "    \n",
    "    return df_benzene\n",
    "\n",
    "if __name__ == \"__main__\":\n",
    "    result_df = extract_all_omi_data()"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "b91cee4b-61ad-4e6f-a708-d5aa553e0fcd",
   "metadata": {},
   "source": [
    "In this project, I built a Python pipeline that enriches benzene observations with OMI satellite data by coding a fast index-based extraction workflow. I first loaded a pre-computed mapping of each monitoring site to its OMI grid indices, then looped through all unique dates in the benzene dataset. For each date, the code locates the corresponding OMI files (NO₂, O₃, SO₂, HCHO), opens them with h5py, and directly slices the arrays at the stored indices instead of doing repeated spatial searches. The script filters out fill values, averages valid orbital measurements, and stores results along with orbit counts and quality flags. Finally, I merged these values back into the benzene dataframe and saved the enriched dataset as a CSV, creating a reproducible coder-friendly workflow that connects satellite fields with ground benzene records."
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3 (ipykernel)",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
   "version": "3.10.13"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}
