Skip to content

API Reference

This page is auto-generated from Python docstrings.

datafun.app

src/datafun/app.py - Project script.

Author: Denise Case Date: 2026-08-23

HOW TO RUN THIS FILE:

From the VS Code menu (with only this project open in VS Code), click "Terminal" / New Terminal to open an integrated Terminal in the root project folder. Paste the following command and press ENTER or RETURN to run this file as a script:

uv run python -m datafun.app

DOMAIN:

This project illustrates how the workflow is similar even when the data is very different. It uses four datasets, each in a different file format.

  • CSV: world happiness scores
  • JSON: astronauts currently in space, by spacecraft
  • XLSX: student feedback text
  • TXT: a plain-text version of Romeo and Juliet

Paths (relative to repo root):

INPUT FILE: data/raw/2020_happiness.csv INPUT FILE: data/raw/astros.json INPUT FILE: data/raw/Feedback.xlsx INPUT FILE: data/raw/romeo_and_juliet.txt

OUTPUT FILE: data/processed/csv_ladder_score_stats.txt OUTPUT FILE: data/processed/json_astronauts_by_craft.txt OUTPUT FILE: data/processed/xlsx_feedback_github_count.txt OUTPUT FILE: data/processed/txt_summary.txt

EXPLORE:

Raw data usually needs work before it can be trusted. An ETVL pipeline moves data through four stages:

  • Extract: read raw values from a source file
  • Transform: calculate results from the raw values
  • Verify: check the results before writing them
  • Load: write the verified results to an output file

The file format and the transform differ for each dataset, but the four ETVL stages remain the same.

DESIGN:

Use this file to declare the data-specific choices and the reasoning behind them, then orchestrate the work. The format-specific ETVL pipelines live in supporting modules.

main

main() -> None

Entry point when running this file as a Python script.

This is where the instructions begin.

Arguments: None. Returns: None.

Source code in src/datafun/app.py
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
def main() -> None:
    """Entry point when running this file as a Python script.

    This is where the instructions begin.

    Arguments: None.
    Returns: None.
    """
    log_header(LOG, "P03")

    LOG.info("===================================")
    LOG.info("START main()")
    LOG.info("===================================")

    log_path(LOG, "raw folder", path=RAW_DIR)
    log_path(LOG, "processed folder", path=PROCESSED_DIR)

    # Ensure the output folder exists before any pipeline writes to it.
    PROCESSED_DIR.mkdir(parents=True, exist_ok=True)

    LOG.info("===================================")
    LOG.info("CSV pipeline")
    LOG.info("===================================")

    LOG.info(f"Pipeline: {CSV_PIPELINE_DESCRIPTION}")
    LOG.info(f"Column: {CSV_COLUMN}")
    LOG.info(f"Why: {WHY_CSV_COLUMN}")
    run_etvl_csv(
        input_file=CSV_INPUT,
        output_file=CSV_OUTPUT,
        column_name=CSV_COLUMN,
        log=LOG,
    )

    LOG.info("===================================")
    LOG.info("JSON pipeline")
    LOG.info("===================================")

    LOG.info(f"Pipeline: {JSON_PIPELINE_DESCRIPTION}")
    LOG.info(f"Grouping: {JSON_LIST_KEY} by {JSON_CRAFT_KEY}")
    LOG.info(f"Why: {WHY_JSON_GROUPING}")
    run_etvl_json(
        input_file=JSON_INPUT,
        output_file=JSON_OUTPUT,
        list_key=JSON_LIST_KEY,
        craft_key=JSON_CRAFT_KEY,
        log=LOG,
    )

    LOG.info("===================================")
    LOG.info("XLSX pipeline")
    LOG.info("===================================")

    LOG.info(f"Pipeline: {XLSX_PIPELINE_DESCRIPTION}")
    LOG.info(f"Word: {XLSX_WORD} in column {XLSX_COLUMN}")
    LOG.info(f"Why: {WHY_XLSX_WORD}")
    run_etvl_xlsx(
        input_file=XLSX_INPUT,
        output_file=XLSX_OUTPUT,
        column_letter=XLSX_COLUMN,
        word=XLSX_WORD,
        log=LOG,
    )

    LOG.info("===================================")
    LOG.info("TEXT pipeline")
    LOG.info("===================================")

    LOG.info(f"Pipeline: {TXT_PIPELINE_DESCRIPTION}")
    LOG.info(f"Why: {WHY_TXT_SUMMARY}")
    run_etvl_text(
        input_file=TXT_INPUT,
        output_file=TXT_OUTPUT,
        log=LOG,
    )

    LOG.info("===================================")
    LOG.info("END main() - Executed successfully!")
    LOG.info("===================================")