[Smokey Interviews] Maëlle Salmon, Statistician with the CHAI project

Hi Maëlle Salmon. Nice to meet you!
Hi Maëlle Salmon. Nice to meet you!
We’re excited to continue this interview series today with another influencer in the air quality world. In case you haven’t met Smokey yet — Smokey is the world’s first and only real-time air quality chatbot. Just send him a message and he will reply in a jiffy with the air quality of your city in Messenger.

Let’s meet Maëlle Salmon, statistician at CHAI project in Barcelona, Spain.

What is the CHAI project? Can you tell us what you do there?

CHAI is a research project lead by Cathryn Tonne about air pollution (PM2.5 and black carbon) and health. The goal of CHAI is to investigate long-term effects of particle exposure on cardio-vascular health. The study is based on the Andhra Pradesh Children and Parents Study (APCAPS) prospective cohort in a rural area south of Hyderabad in Telangana. I work for CHAI from Barcelona, Spain as a statistician and data manager.

What got you interested in air pollution?

I’ve been generally interested in environmental issues for quite a while, but it’s the first time I work in the field of air pollution, so CHAI has quite increased my interest for air pollution issues. Given the high particles concentrations one encounters in CHAI study area, better characterizing their health effect is crucial and currently most studies have been performed in countries with lower particles concentrations.

Why did you create the R-package for OpenAQ?

In CHAI there were three ambient monitoring sites up for one year and as a comparison we use data from the US consulate in Hyderabad. When I started wondering about data from the Indian CPCB (central pollution control board) itself and went to its website I embarked on a kind of scavenger hunt clicking around which felt quite frustrating. This also happens with websites from other countries, with a different scavenger hunt for each website. So when my boss mentioned OpenAQ during a symposium about the future of environmental epidemiology, I got really interested and excited. People making use of public air quality data easier! The coolest thing ever!

Then, as I’m a passionate R user, it seemed quite logical for me to write a R package for accessing the data, since you can then use R cool visualization and modeling capabilities. I first asked Christa whether there was any R package for OpenAQ. I think I’m the main user of ropenaq but I hope it can be useful for anyone wanting to play with the data and I get pretty happy when I hear about other users. The package is now part of rOpenSci which is a project aimed at transforming science through open data. For that ropenaq underwent review at rOpenSci by Andew MacDonald and Andy Teucher but any further bug report or feature request is welcome.

Besides air pollution data, what other datasets have you worked with?

I work in environmental epidemiology, so we’ve got data about health too. Then air quality data goes hand in hand with weather data (see my riem package for accessing weather data from airports). Before working for CHAI I worked in the infectious diseases field with surveillance data. I did a PhD in statistics and public health, about aberration detection in time series of say number of weekly cases of Salmonella (yep like my name). Besides I love playing with data of different sources and a current favorite of mine as a trains and public data lover is this awesome dataset about Indian railways. Given the size of India, putting together all this data about trains, stations and schedules in India must have been a lot of work.

You live in Barcelona where the air is really clean. I’m in New Delhi. Can you remind me what clean air feels like?

Ah, I’m so spoiled that I don’t even know how dirty air feels like… Although Barcelona is not perfectly clean, it’s quite clean in comparison with New Delhi. Let’s say I never look at air quality figures before going out for a run, I’m healthy so the concentrations here don’t give me any acute respiratory symptom. Air quality doesn’t influence my everyday choices at all and I feel very lucky when I compare my situation with yours or the situation of my colleagues in India or that of CHAI participants.

If you weren’t a data scientist, what would you be?

My official job title is statistician, not data scientist. ;-) I love my job so much… I don’t know!

Would you rather fight a horse-sized duck or a 100 duck-sized horses?

Hey why would any of those want to fight me? But well I’d rather fight the horse-sized duck because I hope that Smokey, as a bird, would help me solve the conflict diplomatically with the duck.