DSCI 521 Milestone 3
Reproducible environments and two computational posts
See the assessment table on the home page for the weight, due date, and where to submit.
What you are doing
You are going back to the website from Milestone 2 and making it reproducible.
- Add two computational posts, one Quarto with R, one Quarto with Python.
- Pin one environment per language and commit the lockfiles:
uvfor Python,renvfor R. - Write build instructions in your
README.md. - Clean up the site.
We grade by cloning your repository and following your own build instructions. If the site does not rebuild for us, you lose those marks.
It is week 4. These steps say what has to exist, not which button to click.
Lectures this milestone draws on:
What counts as a computational post
The numbers and figures on the page were produced by code that ran at render time. Not a screenshot of a plot.
Each post needs:
- A dataset.
- A line saying where the data came from, with a link.
- At least three code chunks that do real work.
library()calls do not count. - At least one figure or table your code generated.
- Prose saying what you did and what you found.
- Code and output visible on the rendered page.
There is no upper limit. Write as much as you want.
The two posts can be the same analysis done twice, or two different analyses. Either is fine. What we check is that each post is real work in that language.
Picking a topic and a dataset
We are not giving you a dataset or a question. Finding data, working out what it supports, and telling a story with it is the job. This is a small, cheap place to do that whole loop yourself.
You are marked on whether it runs, whether we can rebuild it, and whether a reader can follow it. Not on how clever the analysis is.
Use any dataset you like
A dataset that ships in a package is the easy route. The licensing is settled, and uv sync or renv::restore() brings the data along with the package.
- R:
palmerpenguins,gapminder, or built-ins likemtcarsandiris. - Python:
palmerpenguins, orscikit-learn’s bundledload_*datasets.
Nothing that needs a login, an API key, or a token. This is a public repository.
seaborn.load_dataset() downloads from the internet on every call. If it needs the network, treat it as a URL, not as a package dataset.
Say where it came from
The source goes in the post itself, with a link. Not only in your README.md or a code comment.
Data: [Palmer Penguins](https://allisonhorst.github.io/palmerpenguins/),
Palmer Station Antarctica LTER.Name the licence if it has one. Cite it the way it asks to be cited.
Check that you are allowed to republish it
Committing a data file to a public repository is publishing it. Free to download does not mean free to redistribute.
Before you commit a data file, find its licence or terms of use and check that redistribution is allowed. CC0, CC BY, and most open government licences allow it, usually with attribution. If you cannot find a licence, assume you do not have permission.
If you cannot republish it, either:
- Pick a different dataset, or
- Read it from its URL at render time instead of committing the file. Your site then only rebuilds while that host is up, so say so in your
README.md.
Keep it small
Committed data files stay under about 5 MB. If your dataset is bigger, commit a sample and say in the post how you sampled it. You wait for every render too.
Where the environment files go
Both environments live at the top level of the repository, next to _quarto.yml. Not inside posts, and not one per post.
username.github.io/
├── _quarto.yml
├── README.md
├── index.qmd
├── about.qmd
├── blog.qmd
├── pyproject.toml <- uv
├── uv.lock <- uv
├── .python-version <- uv
├── renv.lock <- renv
├── .Rprofile <- renv
├── renv/ <- renv
├── posts/
└── docs/
Your data file can sit in the same folder as the post that uses it. Quarto runs a code chunk with the working directory set to that post’s folder, so read_csv("penguins.csv") finds a file next to index.qmd.
Always run quarto render from the top level. That is how R finds .Rprofile and turns renv on.
Step 1: Plan the two posts
Pick your topics and datasets before you install anything. Do the licence check now, while you still have time to change your mind.
Each post gets its own folder, as in Milestone 2:
posts/<short-name>/index.qmd
Lowercase, hyphens, no spaces. See the naming guidelines.
Step 2: Set up the Python environment
Textbook: Creating a project · Adding packages · The two files that matter
From the top level:
uv init --bare
uv python pin 3.14
uv add jupyter ipykernelThen uv add what your post needs, such as pandas or altair.
jupyter and ipykernel are required. Quarto runs Python chunks through a Jupyter kernel, and the kernel has to be in this project’s environment.
Render with uv run so Quarto uses the project’s .venv:
uv run quarto renderTo confirm, print sys.executable in a chunk. The path should end in .venv/bin/python (might be .venv/Scripts/python on Windows) inside your repository. If it does not, you rendered without uv run.
rm -r .quarto
uv run quarto renderStep 3: Set up the R environment
Textbook: R environments · Environment files · Snapshot packages
From an R session started at the top level:
renv::init()That creates .Rprofile, renv/, and renv.lock. Install what your post needs, then:
renv::snapshot()Open renv.lock and confirm your packages are in it. renv finds dependencies by reading your files for library() calls, so a package you only loaded in the console will be missing.
You do not activate anything by hand. Running quarto render from the top level starts R there, which reads .Rprofile and turns renv on.
renv.lock, not your R library
renv writes its own renv/.gitignore that keeps renv/library out of Git. Leave it alone. What belongs on GitHub is renv.lock, .Rprofile, and renv/activate.R.
Step 4: Write the two posts
Each post needs everything in What counts as a computational post.
Give each one a YAML header with title, author, and date, as in Milestone 2. They show up on blog.qmd on their own.
Use chunk options on purpose
Textbook: Code chunk options
Use the Quarto #| style, not inline {r, echo=FALSE}. Across the two posts we want:
A figure with a label and a caption:
#| label: fig-penguin-mass #| fig-cap: "Body mass of penguins by species."A chunk that quiets something on purpose, such as a setup chunk:
#| message: false #| warning: falseYour analysis code visible on the page.
echo: false everywhere hides the thing we are grading. Hide the noise, show the work.
Step 5: Lock what you actually used
You installed something while writing that is not in a lockfile yet.
uv syncrenv::snapshot()Check that uv.lock lists every package your Python post imports, and renv.lock every package your R post loads. One missing package is a site we cannot rebuild.
Step 6: Write the build instructions
Your README.md is a deliverable this week. Write it for somebody who has your repository and nothing else.
Cover:
- What this repository is. One or two sentences.
- What to install first, with the versions you used. Quarto,
uv, and R at minimum.renvbootstraps itself. - The exact commands, in order, from
git cloneto a built site. Copy-pasteable, shell and R, saying where each one is run. - Where the built site lands and how to open it locally.
- Where the data comes from, and whether the build needs the network to fetch it.
Then test it:
# this will clone into another directory and name you specify
git clone git@github.com:username/username.github.io.git ~/tmp/m3-test
cd ~/tmp/m3-testFollow your own README line by line, without fixing anything from memory. Anything you have to guess at is a line missing from your README.
Step 7: Clean up the site
On the live site, not your local preview:
- Every navbar link works.
- Every image loads.
- The blog listing shows all your posts, newest first.
- No broken internal links.
- No leftover template text, such as Quarto’s starter “About this site”.
- Your home and about pages still read well.
On github.com:
docs/.nojekyllis there..quarto,_site,.DS_Store,.venv, andrenv/libraryare not.- The repository is still Public.
You learned ignore files in Week 3. Add a .gitignore at the top level with .quarto/, _site/, .DS_Store, and .venv/. Textbook: Create a .gitignore file.
If .venv is already on GitHub, adding it to .gitignore changes nothing:
git rm -r --cached .venv
git commit -m "stop tracking the virtual environment"
git push origin mainStep 8: Render, commit, push, and check
From the top level:
uv run quarto render
git status
git add .
git commit -m "milestone 3"
git push origin mainWait for GitHub to build, then open https://username.github.io in a private window and click through both posts. Both render, code and output are visible, figures appear with their captions.
Commit as you go, not once at the end. We expect at least five commits.
Step 9: Submit
A PDF with two URLs:
https://github.com/username/username.github.iohttps://username.github.io
Upload it to Gradescope. The bonus adds two more things to that PDF. See Submitting the bonus.
Bonus: R and Python in one post
Worth 5 marks.
Write a third post that runs R and Python in the same document and passes an object between them. One page, both languages, a result computed in one and used in the other.
You choose where the 5 marks go: this milestone, or any one other assignment in this course.
- They are not split. All five land on one assignment.
- No assignment goes above 100%, so spending them on something you already aced wastes them.
- Tell us where to put them when you submit. If you do not say, we apply them here.
The mechanism is reticulate. A .qmd with {r} chunks uses the knitr engine, and knitr hands {python} chunks to reticulate in the same session.
What you need to work out:
- Getting
reticulateintorenv.lock. - Pointing
reticulateat this project’s.venv. Tryuv run quarto renderfirst, and set it explicitly if that picks the wrong Python. py$namereads a Python object from R, andr.namereads an R object from Python.
References:
- Quarto: Using Python, the knitr engine
- reticulate: Calling Python from R
- reticulate: Python version configuration
The bonus counts only if the post renders when we follow your README from a clean clone.
Submitting the bonus
We do not go hunting for it. Add both of these to your Gradescope PDF:
- A separate URL that opens the bonus post itself, such as
https://username.github.io/posts/r-and-python/. Not the repository, not the home page. - One line naming the assignment the 5 marks go to, for example “Apply the bonus to Milestone 2.”
Before you submit, check
If you did the bonus post:
Grading
| Item | Marks |
|---|---|
| Python post: real analysis, code and output on the live site, data source documented | 7 |
| R post: real analysis, code and output on the live site, data source documented | 7 |
| Chunk options used on purpose in both posts | 3 |
Python environment: pyproject.toml, uv.lock, .python-version, complete |
4 |
R environment: renv.lock, .Rprofile, renv/activate.R, complete |
4 |
README.md build instructions we can follow from a clean clone |
6 |
| Site cleanup: navigation, links, images, no leftover placeholder text | 2 |
| Site live on GitHub Pages, at least five commits, both URLs on Gradescope | 2 |
| Total | 35 |
| Bonus | Marks |
|---|---|
| A third post that runs R and Python together and passes an object between them | +5 |
The 5 bonus marks go to this milestone or to any one other assignment in this course, whichever you name when you submit. They are not split, and no assignment goes above 100%.
If you get stuck
Textbook: Asking Effective Questions
Ask in the 521_platforms-dsci Slack channel or come to office hours. Post the command you ran and the error message you got, not just “it did not work”.