Home > Net >  dplyr summarise : Group by multiple variables in a loop and add results in the same dataframe
dplyr summarise : Group by multiple variables in a loop and add results in the same dataframe

Time:11-14

I want to calculate indicators on the different modalities of several variables, and then add these results in a single dataframe. I can do this without any problem with several summarise coupled with group_by, and then do a rbind to gather the results. Below, I do it on the hdv2003 data (from questionr package), and I rbind results created on variable 'sexe', 'trav.satisf' and 'cuisine'.

library(questionr)
library(tidyverse)
data(hdv2003)

tmp_sexe <- hdv2003 %>%
  group_by(sexe) %>%  
  summarise(n = n(),
            percent = round((n()/nrow(hdv2003))*100, digits = 1),
            femmes = round((sum(sexe == "Femme", na.rm = TRUE)/sum(!is.na(sexe)))*100, digits = 1),
            age = round(mean(age, na.rm = TRUE), digits = 1)
  )

names(tmp_sexe)[1] <- "group"

tmp_trav.satisf <- hdv2003 %>%
  group_by(trav.satisf) %>%  
  summarise(n = n(),
            percent = round((n()/nrow(hdv2003))*100, digits = 1),
            femmes = round((sum(sexe == "Femme", na.rm = TRUE)/sum(!is.na(sexe)))*100, digits = 1),
            age = round(mean(age, na.rm = TRUE), digits = 1)
  )

names(tmp_trav.satisf)[1] <- "group"

tmp_cuisine <- hdv2003 %>%
  group_by(cuisine) %>%  
  summarise(n = n(),
            percent = round((n()/nrow(hdv2003))*100, digits = 1),
            femmes = round((sum(sexe == "Femme", na.rm = TRUE)/sum(!is.na(sexe)))*100, digits = 1),
            age = round(mean(age, na.rm = TRUE), digits = 1)
  )

names(tmp_cuisine)[1] <- "group"

synthese <- rbind (tmp_sexe,
                   tmp_trav.satisf,
                   tmp_cuisine)

Here's the result :

# A tibble: 8 x 5
  group              n percent femmes   age
  <fct>          <int>   <dbl>  <dbl> <dbl>
1 Homme            899    45      0    48.2
2 Femme           1101    55    100    48.2
3 Satisfaction     480    24     51.5  41.4
4 Insatisfaction   117     5.9   47.9  40.3
5 Equilibre        451    22.6   49.9  40.9
6 NA               952    47.6   60.2  56  
7 Non             1119    56     43.8  50.1
8 Oui              881    44     69.4  45.6

The problem is that this writing is too long and not manageable. So I would like to produce the same result with a for loop. But I have a lot of trouble with loop in R and I can't do it. Here's my try :

groups <- c("sexe",
            "trav.satisf",
            "cuisine")

synthese <- tibble()

for (i in seq_along(groups)) {
  tmp <- hdv2003 %>%
    group_by(!!groups[i]) %>%  
    summarise(n = n(),
              percent = round((n()/nrow(hdv2003))*100, digits = 1),
              femmes = round((sum(sexe == "Femme", na.rm = TRUE)/sum(!is.na(sexe)))*100, digits = 1),
              age = round(mean(age, na.rm = TRUE), digits = 1)
    )
  
  names(tmp)[1] <- "group"
  synthese <- bind_rows(synthese, tmp)
}

It works but it doesn't produce the expected result, and I dont understand why :

# A tibble: 3 x 5
  group           n percent femmes   age
  <chr>       <int>   <dbl>  <dbl> <dbl>
1 sexe         2000     100     55  48.2
2 trav.satisf  2000     100     55  48.2
3 cuisine      2000     100     55  48.2

CodePudding user response:

library(questionr)
library(tidyverse)
data(hdv2003)

list("trav.satisf", "cuisine", "sexe") %>%
  map(~ {
    hdv2003 %>%
      group_by_at(.x) %>%
      summarise(
        n = n(),
        percent = round((n() / nrow(hdv2003)) * 100, digits = 1),
        femmes = round((sum(sexe == "Femme", na.rm = TRUE) / sum(!is.na(sexe))) * 100, digits = 1),
        age = round(mean(age, na.rm = TRUE), digits = 1)
      ) %>%
      rename_at(1, ~"group") %>%
      mutate(grouping = .x)
  }) %>%
  bind_rows() %>%
  select(grouping, group, everything())
#> # A tibble: 8 x 6
#>   grouping    group              n percent femmes   age
#>   <chr>       <fct>          <int>   <dbl>  <dbl> <dbl>
#> 1 trav.satisf Satisfaction     480    24     51.5  41.4
#> 2 trav.satisf Insatisfaction   117     5.9   47.9  40.3
#> 3 trav.satisf Equilibre        451    22.6   49.9  40.9
#> 4 trav.satisf <NA>             952    47.6   60.2  56  
#> 5 cuisine     Non             1119    56     43.8  50.1
#> 6 cuisine     Oui              881    44     69.4  45.6
#> 7 sexe        Homme            899    45      0    48.2
#> 8 sexe        Femme           1101    55    100    48.2

Created on 2021-11-12 by the reprex package (v2.0.1)

  • Related