Showing posts with label R. Show all posts
Showing posts with label R. Show all posts

Tuesday, 13 August 2013

Social Constructivism in Big Classes

+Theron Hitchman (who is a great guy to follow on G+ and one of the go to guys as far as I am concerned when it comes to Inquiry Based Learning (IBL)) recently posted on G+ wondering how he was going to manage with a course of his having 68 students enrolled.

I am still understanding IBL but it is one of many teaching methodologies that use a constructivist model (here's a post by +Dana Ernst about an article that claims that these approaches aren't that great). I have spoken about constructivism before when reviewing a great education book. The basic concept of constructivism is that students need/should learn through a personal 'construction of the concepts'. IBL does this by giving students the reigns to the direction and also the discovery of the content of a course (I am sure +Dana Ernst+Theron Hitchman  and others could either forgive me or correct my brief summary).

This is often done (although not exclusively, take a look at +Bret Benesh's post) done through the use of student presentations and generally needs a non-lecture based approach: which is difficult with a large number of students!.

Theron's post on G+ got a lot of comments with people offering advice and +Bret Benesh wrote a nice blog post suggesting a peer learning approach.

I thought I would throw in my two cents as I have recently designed a new course that made use of peer learning.

The course is a statistical programming course on our MSc programme here at +Cardiff University.

You can see all the resources for the course here.

To deal with large teaching numbers in that class (this was the first year of running and it had 27 students but I certainly believe that my approach is scalable and I designed it thinking that I'd have 40+ students at times). I used a peer learning approach based on flipped classrooms, problems and with a hint of IBL thrown in.
  • First of all it took a lot of preparation. I flipped the class completely putting together a full set of notes as well as a full set of videos (you can see them on the site).
  • I made groups of 3/4 out of all the students.
  • For every topic, I put together a set of problems that all the groups where told they had to complete. I called these 'challenges' (it seemed a bit snappier) and you can see them on the course website.
  • Every class (this course was taught in 8 sessions over a 4 week period) started with the me asking who would like to present a solution to the challenge.
The students were aware that their final mark for the course would take in to account participation in these presentations but I made it clear that the corresponding contribution would be mainly based on the fact that they tried and not necessarily how good a job they did. In essence I did not really care whether or not they managed to do the challenges, I just wanted them to work on it!

By the beginning of the 5th session, students were happy to present and needed no gentle nudging from me. In fact some of the students who presented started off by saying "I didn't manage to do this but this is what I tried". Seeing that students got that it was ok to not succeed at something was really great to see. Whenever that happened the students would start discussing amongst themselves how to do a particular problem: and would teach each other!

Throughout all of this, I would sit at the back of the room and vary rarely say anything at all. At times I would perhaps nudge a bit with a "does anyone know a different way to do that?".

One of my favourite memories of the course was towards the end when at the beginning of one session I did not prompt the students to start the class. All of them were just chatting amongst themselves (not about the course) and I was sat at the back observing. After a while one of the students turned around and looked at me questionably. I just smiled. She got up, told everyone that they should get started and the class started teaching itself. That day I didn't actually say anything.

I think this happened to work well because the challenges were designed in such a way so that students were lead through the curriculum but given the nature of the course I also made sure that they had plenty of resources with which they could practice things further.

I encouraged the students to find other resources then the ones I put together and also encouraged them to share their solutions with each other after each session (I believe they used a dropbox folder but I emphasised that I would stay out of the way).

The whole course culminated with one of the assessments being a presentation where the students (still in their group) had to teach me something that was not in the resources made available to them. I wrote about this here and one of the things a group put together was this cool gif of the Mandlebrot set:


I'm currently going through a teaching certification process and have blogged about it many times. You can see my various portfolios [here](http://www.vincent-knight.com/home/teaching/pcutl) but if I am to take only one sentence from everything I've done so far it would be this pretty cool definition of teaching:

"To create learning opportunities."

I believe that teaching in a constructivist framework does exactly that and I really hope that large classrooms do not dissuade educators from using it :)

Anyway, what I hoped would be a perhaps helpful post for +Theron Hitchman ended up just being another page or two of me rambling on...

Saturday, 8 June 2013

Student choices between SAS and R in Individual Coursework

This is the third post is a small series of posts reflecting on my teaching in a new class (all teaching materials can be found here) introducing students to SAS and R on our MSc course. The two previous posts have so far just been a reflection on student's attitudes towards each piece of software.
In the first of those posts I said how I was slightly surprised at how students had chosen to use SAS for 1 particular question where they had the option of the language. In my opinion it was a problem much easier to tackle with R (all the course works, class tests etc can be found on this site). I also mentioned in that post how I asked students which language they preferred. Almost all students answered that it depends on the task (which is a great answer) but after pushing them for a particular decision a strong majority seemed to prefer R. I posted on G+ recently about a particular interaction I've had subsequently with a student which seemed to confirm this attitude of needing to find the correct tool for the correct job.

In the second post I described how students in their group presentations (asking them to teach me something I had not taught them) mostly evaluated SAS v R for certain tasks. It was great to see them identify strengths and weaknesses for each language.

This post is about their choices in their individual coursework.

Similar to their class test (which I discuss in the first post of this series) there was a question which allowed students to choose a language in this individual coursework component of the class (which can be found here).

In my opinion this question made use of quite a big data set (generated by some research I'm currently doing) and I thought it would probably be simpler to approach in SAS. About 52% of the class agreed with me whilst 48% seemed to still prefer R. First of all, I could be wrong and R could indeed be better suited for this question, secondly it might also be a reflection of the personal preferences that the students seemed to indicate when I asked: most students seemed to prefer R. If the latter is the case then I suppose it's nice to see that students not just realise that there's a better tool for a given job but also a better tool for a particular person doing a given job. I'll be keeping a track of this over the years and see how (if) it changes.

In my next post in this series I'll start to reflect on some of the teaching methodologies I used.

Saturday, 25 May 2013

Student choices between SAS and R in teaching presentations.

This is my second post in a series of posts (the first one is here) about a SAS/R course that I've recently finished teaching to MSc students at +Cardiff University.

I taught this course using a hybrid of flipped classrooms / IBL (although I'm cautious when using the term IBL as I'm not entirely sure my approach fits with any variation of Moore's methods). I gave students access to all the content of the class before hand (including notes, exercises and a series of screencasts - all the materials are here if they're of interest). The students were then given "challenges" and had to deliver their solutions to as presentations to the other students in the class. The aim of this was to get the students to teach themselves/each other and quite often I would not actually have to say much at all during a class (this allowed for a better use of 'me' by the students during the lab sessions).

(Here's a previous post about flipping the classroom and here's one about IBL)

In the previous post I described how students chose to use SAS and/or R in their class test. Most students chose SAS despite displaying a preference for R when asked.

The above is 1 of 3 assessments that the students have had to go through. This post is about the second assessment: a group presentation. In this presentation I asked students to teach an aspect of SAS or R that had not been covered in class (you can see the brief here).

I believe that this is a particularly important thing to assess as I in no way can pretend to teach them everything. It's important that they know how to learn new things that they might need in their career.

I was expecting groups to select a particular language and then a particular topic but interestingly most groups chose to look at both languages and compare strengths and merits.
I had 6 groups and here are the subjects that they looked at:
  • Time series forecasting: both in SAS and R;
  • Principal component analysis: both in SAS and R;
  • Random sampling: both in SAS and R;
  • Survey sampling (in SAS) and the creating a gif of the Mandelbrot set (in R);
  • Scorecard building in SAS;
  • Mapping and spatial analysis: in R.
The 3 groups who carried out a single thing in both SAS and R did a good job of describing strengths and weaknesses of both languages. It was a pleasure to see and again reassures me that an important message has gotten across to the students which is that there is not 1 best tool but an appropriate tool for a particular job.

I'm planning on putting the code/slide up for these talks for them to serve as resources for students doing the course next year but I want to wait till I've marked their final piece of work and asked the students if they mind. In the mean time I'll repost this gif of the Mandelbrot set made by 1 of the groups (I thought this was cool!):



I still want to write generally about the teaching/learning methods in this class and will do that later but if it's of interest my PCUTL portfolio is available here and in there I describe and justify a lot of what I'm doing.

Saturday, 18 May 2013

Student choices between SAS and R

I'm going to be writing a couple of posts looking back at a class I've taught that's just coming to an end (at the time of writing this I've got one more group presentation to see).

The course teaches SAS and R in parallel on our MSc course (if it's of interest all the teaching materials are here).

I'll be blogging about this class as I taught it a bit differently to the usual "Students Listen - Teachers Lecture" style. I'll get back to that more in future posts (although a lot of what I've done is in my PCUTL module 2 portfolio).

The purpose of this post will be to briefly discuss two questions that were on my class test that I feel give some (very shallow) information as to how the students experienced the course. (The class test was made of 4 questions: q1 - a simple task to be performed in both SAS and R, q2 - a task in SAS, q3 - a task in R and finally q4 - a task in either language.)

The class is taught over 5 weeks:
  1. Week 1: Introduction and basic statistics
  2. Week 2: Data manipulation
  3. Week 3: Programming
  4. Week 4: Extras (for example we take a look at proc optmodel and ggplot2).
  5. Week 5: A 2 hour class test
The first question on the class test (you can see it here) asked the students to rank their enjoyment of each week (the purpose of this question was to give them a nice easy starting point). So a ranking of 1 implied a favorite week while a ranking of 5 implied the least favorite.

This plot shows the mean ranking given to each week:


This data and the following discussion should be taken with no implied rigour: I'm not analysing this too closely and also students might just have written down any sequence of rankings without thinking about things too much (this was after all a test).


First of all it does seem that the students enjoyed the class test the least (which I guess is to be expected).

Secondly, it looks like the first week was perhaps less enjoyed in general.

I think this is also to be expected, I taught the class in a way that I don't believe would have been familiar to the students (I tried to encourage them to teach each other and themselves) so perhaps that first week was just a bit too unfamiliar. I'll try and rectify that in future years (if only by pointing students to what students from previous years did).

Now for the second point.



By design I hope that the students learn how to carry out various programming tasks in SAS and R seeing the strengths and weaknesses of each language as they go. The last question of the class test (again here) involved a bit of data manipulation on small data sets and I believe that the main difficulty was that this question did not force a language on the students. In essence choosing the language was the most important point of the question.

My personal approach would have been to use R for this particular question but interestingly most students chose SAS. Some used a combination (often starting in SAS before realising that perhaps R was better suited which led to a bit of a clumsy hybrid). On average SAS was used by 77% of the students for question 4 (some of which managed the task very well!).

During the group presentations I've been asking students afterwards a "this does not count" question (ie making it clear that there's no wrong or right answer to this and that I'm simply interested/curious):
'If you were starting a consultancy company tomorrow but could use only one package: ether SAS or R which would you pick?'
The really pleasing thing is that almost all of the students miss the constraint in my question and immediately reply something like:
"It depends on the kind of consultancy we'll be doing."
I re-iterate the constraint (after telling them that that's the actual right answer :)), I'd say that a majority of students seem to prefer R. Perhaps the biais towards SAS in the class test was due to the conditions (time was short) but overall it's been nice to see that most students realise that it's about finding the right tool for a particular job.

I'm yet to see the individual course work that they'll be handing in this week, which is of a very similar format. I wonder which language they'll have picked for Q4...

EDIT: Here's the next post in this series (looking at choices between SAS and R during teaching presentations).

Sunday, 30 December 2012

Some analysis of the game shut the box

This Christmas, +Zoe Prytherch and I got my mother a small board game called: "Shut the Box" (also known as "Tric-Trac", "Canoga", "Klackers", "Zoltan Box", "Batten Down the Hatches", or "High Rollers" according to the wiki page). It's a neat little game with all sorts of underlying mathematics.
Here's a picture of us playing:


Here's a close up of the box:


A good description of the game can be found on the wiki page but basically the game can be described as follows:
  • It is a Solitaire game that can be played in turns and scores compared (but no strategies arise due to interactions of players).
  • The aim is to "shut" as many tiles (each being one of the digits from 1 to 9) as possible.
  • On any given go, a player rolls two dice and/or has a choice of rolling one dice if the tiles 7,8,9 are "shut".
  • The sum total of the dice roll indicates the sum total of tiles that must be "shut". If I roll 11, I could put down [2,9],[3,8],[1,2,8] etc...
  • The game ends when a player can't "shut" tiles corresponding to the dice roll (including the case when all tiles are "shut").
Here's a plot of a game we played with a few of my mother's friends:



You can see that myself and Jean did fairly poorly while Zoƫ made the difference just at the end to win :)

The game could be modelled as a Markov decision process but over the past few days I decided to code it in Sage and investigate whether or not I could find out anything cool with regards to strategies. If you're not aware of Sage I thoroughly recommend you taking a look, it's an awesome open source mathematics package. In this instance I used the Partitions function to rapidly obtain various combinations of dice rolls and/or tile options to get the results I wanted.

The code is all in a github repo and I feel I've written an ok README file explaining how it works so if you're interested in playing with it please do! The code basically has two parts, the first allows you to play using the program instead of a game board (with prompts telling you what your available plays are).

Here's a quick screencast of me demonstrating some of the code:



The second part of the code allows for strategies to be written that play the game automatically ("autoplay"). I've considered 4 strateges:
  • Random (Just randomly pick any potential tile combination)
  • Shortest (Choose the tile combination that has shortest length: i.e. [3] instead of [1,2] ~ in case of a tie pick randomly)
  • Longest (Choose the tile combination that has longest length: i.e. [1,2] instead of [3] ~ in case of a tie pick randomly)
  • Greedy. This chooses the tile combination that ensures the best possible chance of the game not ending on the next go. This is calculated by summing the probabilities of obtaining a dice roll that could be obtained.
I've been leaving my computer run the code in the background for a while now so here are some early results. I've got over 90000 instances for Greedy and Random with over 40000 instances for the other two which I thought of later (EDIT: I know have over 400000 instances...). I've collected (as and when one of my home boxes was ideal) more data since writing this and the file is available for anyone to play with here. If I've made any mistakes with my analysis, I'd love to hear it!.

First of all what does the average score look like:

Greedy_Score   Random_Score     Longest_Score   Shortest_Score  
Min.   : 0.00  Min.   : 0.00    Min.   : 0.00   Min.   : 0.00  
1st Qu.: 6.00  1st Qu.:14.00    1st Qu.:18.00   1st Qu.: 7.00   
Median :11.00  Median :21.00    Median :24.00   Median :11.00
Mean   :11.42  Mean   :20.45    Mean   :24.16   Mean   :12.06  
3rd Qu.:16.00  3rd Qu.:27.00    3rd Qu.:30.00   3rd Qu.:17.00  
Max.   :43.00  Max.   :43.00    Max.   :43.00   Max.   :43.00  

Shortest and Greedy seem to do a fair bit better then the other two. If we take a look at the distribution of the scores that is confirmed:




If we also take a look at the length of the game (i.e. how many total dice rolls there were):

Greedy_Length  Random_Length  Longest_Length  Shortest_Length
Min.   :1.000  Min.   :1.000  Min.   :1.00    Min.   :1.00    
1st Qu.:4.000  1st Qu.:3.000  1st Qu.:2.00    1st Qu.:4.00    
Median :5.000  Median :3.000  Median :3.00    Median :5.00    
Mean   :4.897  Mean   :3.414  Mean   :2.86    Mean   :4.82    
3rd Qu.:6.000  3rd Qu.:4.000  3rd Qu.:4.00    3rd Qu.:6.00    
Max.   :9.000  Max.   :9.000  Max.   :8.00    Max.   :9.00    
                                                             

We see that games last longer when using Greedy and Shortest. If we look at the distribution we see (I've updated this graph since first publishing this post):



It looks like Greedy might allow for slightly longer games than Shortest. Whether or not playing Shortest or Greedy is actually any different seems interesting. A simple analysis of variance (ANOVA) seems to indicate that there is a significant effect:
Df Sum Sq Mean Sq F value P
stbaov$Method 1 11337 11337 196.3 2e-16
Residuals 127169 7345725 58

This is all experimental of course and analysing whether or not my suggested greedy strategy is indeed the optimal strategy would require some markov decision process modelling. I might look in to this with an undegraduate project student next year.

The tricky part about using the greedy method whilst actually playing the game with friends is that it requires a fair bit of calculation. Nothing that Sage can't handle in a split second but something that could keep people waiting a fair bit if you were to work it out with a pen and paper. Despite the significant statistical difference between the methods this analysis seems to indicate that choosing the shortest set of tiles is a pretty good strategy so if in doubt I recommend choosing that...

If anyone cares enough about all this to fork the code or suggest a different strategy I'd be delighted :) In case you missed it above, here's the github repo:

https://github.com/drvinceknight/Shut_The_Box

(On a technical note some of the pre analysis is done using python and the actual graphs seen here are done using R)