Link To TurkPrime

How to run successful experiments and get the most out of Amazon's Mechanical Turk

Showing posts with label mturk. Show all posts
Showing posts with label mturk. Show all posts

Wednesday, August 15, 2018

TurkPrime Tools to Help Combat Responses from Suspicious Geolocations

Last week, the research community was struck with concern that “bots” were contaminating data collection on Amazon’s Mechanical Turk (MTurk). We wrote about the issue and conducted our own preliminary investigation into the problem using the TurkPrime database. In this blog, we introduce two new tools TurkPrime is launching to help researchers combat suspicious activity on MTurk and reiterate some of the important takeaways from this conversation so far.

TurkPrime’s Tools to Deal with Suspicious Activity

As we announced last week, we’ve created two new tools to help researchers fight fraud in their data collection:

1.      Block Suspicious Geolocations
2.      Block Duplicate Geolocations

The Block Suspicious Geolocations tool is a Free Feature that allows researchers to block submissions from a list of suspicious geolocations. In our investigation last week, we identified several geolocations that were responsible for a majority of duplicate submissions. Our Block Suspicious Geolocations tool will prevent any MTurk Worker from submitting a HIT from these locations. As mentioned in last week’s blog, once we removed these locations from our analyses, we saw the rate of duplicate submissions from the same geolocation across studies this summer fell to 1.7%—a number well within the range of what we’ve identified as normal across the life of our platform. The screenshot below shows our new Block Suspicious Geolocations tool, found in Tab 6 “Worker Requirements” when you design a study.

Our second tool, the Block Duplicate Geolocations tool, is a Pro Feature that allows researchers to block multiple submissions from any geolocation. The Block Duplicate Geolocations tool casts a much wider net than the Block Suspicious Geolocations tool and should ensure that responses collected in any one survey come from a more distributed set of locations. By restricting the number of submissions from each geolocation, researchers can be more confident that the responses they collect are coming from unique participants. When using this tool data collection may be a little slower, especially if the target sample is  concentrated in a small geographic area (e.g., one particular state). The screenshot below shows our new Block Duplicate Geolocations Tool, found in Tab 8 “Pro Features” when you design a study.


Moving Forward

Understanding what has caused the recent increase in low quality responses on MTurk and the corresponding increase in submissions from the same geolocation is a matter of ongoing research. As we learn more details we will share them with the research community and continue to develop tools that ensure the highest quality of research data.

More immediately, we have identified a list of worker IDs that have repeatedly been associated with suspicious geolocations. In addition to the tools described above, we will create an internal exclusion list based on the worker IDs of suspicious accounts over the next several days. This exclusion list will create an additional layer of protection on our system by blocking worker accounts that have a high likelihood of being involved in fraud. We will write another blog to provide more detail about this issue in the coming days. In the meantime, however, researchers already have two powerful tools for eliminating fraud in their data collection. These tools should increase researchers’ confidence that they are obtaining genuine responses from unique workers.


Friday, August 10, 2018

Concerns about Bots on Mechanical Turk: Problems and Solutions


Data quality on online platforms
When researchers collect data online, it’s natural to be concerned about data quality. Participants aren’t in the lab, so researchers can’t see who is taking their survey, what those participants are doing while answering questions, or whether participants are who they say they are. Not knowing is unsettling.

Recently, the research community has been consumed with concern that workers on Amazon’s Mechanical Turk (MTurk) are cheating requesters by faking their location or using “bots” to submit surveys. These concerns originated and have been driven by reports from researchers that there are more nonsensical and low-quality responses in recent studies conducted on MTurk. In these studies, researchers have noticed that several low-quality responses are pinned to the same geolocation. In this blog, we’d like to add some context to the conversation, share the findings from our internal inquiry, and inform researchers what TurkPrime is doing to address the issue.

Concern about Bots
The recent concern about bots appears to have begun on Tuesday, August 7th, 2018, when a researcher asked the PscyhMap Facebook group if anyone had experienced an increase in low quality data. In just the third response to that thread, another researcher suggested, “maybe a machine?” Soon, other researchers were reporting an increase in nonsense responses and low-quality data, although at least a few reported no increase in junk responses to their studies. The primary piece of evidence causing researchers to suspect bots was that most of the low-quality responses were tagged to the same geolocation and a few places in particular—Niagara Square in Buffalo, NY; a lake in Kansas; and a forest in Venezuela. What’s more, many respondents from these geolocations provided suspicious responses to open-ended questions, often answering with “GOOD STUDY,” or “NICE.”

Although this activity raises concerns, the conversation, so far, has overlooked some important points. Most critically, while it is clear some researchers have unfortunately collected several bad responses, the research community does not yet know how widespread this problem is. Diagnosing the issue requires knowing how many studies don’t fit the pattern, as well as how many do.

Scope of problem
At TurkPrime, we track the geolocation of all surveys submitted in studies run on our platform. In the last 24 hours, we have worked to determine whether there is a growing problem of multiple submissions from the same geolocation. In reviewing over 100,000 studies that have been launched on TurkPrime, we see that the rate of submissions from duplicate geolocations typically bounced from less than 1% to 2.5% within a study—a number that could be explained by people submitting surveys from the same building, office, internet service provider, or even the same city. Geolocations are not precise, an issue we will discuss in more detail in a future blog post.

Based on this analysis, we set 2.5% as the threshold for detecting suspicious activity. Over 97% of studies have not reached this threshold, showing that the overwhelming majority of studies have not been affected by data coming from  the same geolocation. 

However, when we look at the rate of duplicate submissions based on geolocation over time, we see that in March of this year the percentage of duplicate submissions began edging up. Clearly, this is a problem, but a problem that has emerged only recently.

What TurkPrime is Doing
At TurkPrime, we are developing tools that will help researchers combat suspicious activity. We have identified that all suspicious activity is coming from a relatively small number of sources. We have additionally confirmed that blocking those sources completely eliminates the problem. In fact, once the suspicious locations were removed, we saw that the number of duplicate submissions had actually dropped over the summer to a rate of just 1.7% in July 2018.

What you can do to eliminate the problem
In the coming days, we will launch a Free feature that allows researchers to block suspicious geolocations. This means researchers will be able to block workers from suspicious geolocations, excluding submissions from those locations in their data collection. We will also launch a Pro feature that allows researchers to block multiple submissions from the same geolocation within a study. This feature will cast a wider net and may block well-intentioned workers using the same internet service provider, or working in the same library. This tool will give researchers greater confidence that they are not receiving submissions from anyone using the same location to submit junk responses.

Conclusions
Our data, and the work of multiple researchers, show there has been a recent increase in the number of low quality responses submitted on Mechanical Turk. Data from the TurkPrime database show that the vast majority of all studies, and the vast majority of recent studies, have never been affected by the current concern of bots. What we still don’t know about the recent issue, is whether these responses are coming from bots or foreign workers using a VPN to disguise their location and submit surveys intended to sample US workers. Either way, in the coming days TurkPrime will release tools that allow researchers to block workers from suspicious locations and to decide how narrowly they would like to set the exclusion criteria. Concerns about bots and low quality data on MTurk are not new. But at TurkPrime we will continue to look for ways to ensure quality data and to make conducting online research easier for researchers.     

Thursday, January 11, 2018

MicroBatch is now a Pro Feature

TurkPrime is announcing a change in our pricing for the MicroBatch feature. MicroBatch is now included as a Pro feature, with a fee of 2 cents + 5% per complete. This will also provide users with access to all other pro features, with no additional charge. This change is necessary so that we can continue to provide the highest quality service and tools that our users expect.


To explain the new pricing structure further, see the following example (the part highlighted in blue reflects the recent change):


Study Example: # of HITs: 100
Payment per HIT: $1


Without MicroBatch:
Total Study Cost = $140
With MicroBatch:
Total Study Cost = $127
Payment to Workers: $100
Payment to Workers = $100
MTurk Fees (40%) = $40
MTurk Fees (20%) = $20
TurkPrime Fees (0%) = $0
TurkPrime Fees (2 cents + 5% per complete) = $7

Savings of 13%

Since its inception 3 years ago, TurkPrime has been dedicated to enhancing the quality and usability of online research with the goal of empowering researchers. The Team at TurkPrime remains committed to this mission. As more researchers come to use our platform and the online sampling world continues to evolve, we have taken on more staff, so that we can continue to provide individualized customer support to each of our users, and develop new tools to expand and improve functionality.

We greatly appreciate our relationship with our users, and know that this change will allow us to continue providing high quality service and equipping the research community with the tools and features they need. We also know that for some users this change was unexpected.  If you have a funded grant with funds allocated to participant collection based on our previous pricing, please contact us so that we can honor our previous pricing. We look forward to your feedback, and to assisting you in getting the most out of your research and TurkPrime.

Friday, December 15, 2017

New Feature: Exclude Highly Active Workers

Some workers on MTurk are extremely active, and take the majority of posted HITs. This can lead to many issues, some of which are outlined in our previous post. Although MTurk has over 100,000 workers who take surveys each year, and around 25,000 who take surveys each month, you are much more likely to recruit highly active workers who take a majority of HITs. About 1,000 workers (1% of workers) take 21% of the HITs. About 10,000 workers (10% of workers) take 74% of all HITs.

TurkPrime now has a feature to allow researchers to exclude the most active workers so that you can collect data from less experienced workers who are less likely to have previously taken part in research similar to your own. Below is a screenshot of the “Naivete (Exclude most active Workers)” feature. You can select what percentage of workers you would like to exclude from the dropdown menu seen below. 

Friday, December 8, 2017

Best recruitment practices: working with issues of non-naivete on MTurk

It is important to consider how many highly experienced workers there are on Mechanical Turk. As discussed in previous posts, there is a population pool of active workers in the thousands, but this is far from exhaustible. A small group of workers take a very large number of HITs posted to MTurk, and these workers are very experienced and have seen measures commonly used in the social and behavioral sciences. Research has shown that when participants are repeatedly exposed to the same measures, this can have negative effects on data collection, changing the way workers perform, creating treatment effects, giving participants insight into the purpose of some studies, and in some cases impact effect sizes of experimental manipulations. This issue is referred to as non-naivete (Chandler, 2014; Chandler, 2016).

The current standard approaches to recruitment on MTurk actually compound this problem. When recruiting workers on Mechanical Turk, requesters have the ability to selectively recruit workers based on specific criteria such as the number of HITs previously approved, and the workers’ approval rating - the percentage of previous HITs that were approved of total HITs completed. A commonly used standard is to select workers who have approval ratings of >95% (see Peer, 2014). This is not quite enough on its own, however, because MTurk’s system assigns a 100% approval rating to all workers who have completed between 1 and 100 HITs, regardless of how many were actually approved. Once workers complete 100 HITs, their approval rating accurately reflects the number of HITs they were approved for. It is therefore recommended, and common practice to only recruit workers who have approval ratings of >95% and who have completed at least 100 HITs. Once researchers use the approval rating system as part of their qualifications for a study, by default, the TurkPrime system adds the qualification that workers must have previously completed at least 100 HITs in order to address this issue (researchers do of course have manual control of this).

By selectively recruiting workers with a high approval rating and a high number of previously completed HITs a requester can have increased confidence that workers in their sample can be trusted to follow instructions and pay attention to tasks. Indeed, many researchers choose to recruit participants who have high approval ratings and have completed a high number of previous studies. The approval rating system is unique, as it is a constant motivating factor that makes workers pay attention to each task. This system helps researchers collect high quality data. However, this leads to the exclusion of workers who have completed few HITs, even if they may be good providers of data but haven’t yet had the chance to “prove” it. The use of only workers who have high approval ratings has a negative effect, which is that it is a selection criteria that is based on recruiting only workers who are more experienced, and therefore less naive to measures used on the MTurk platform, bringing the issue of non-naivete to the fore.

Solutions
TurkPrime is introducing a new tool which allows requesters to exclude workers who are extremely active, thus making it possible to selectively recruit workers who are not overly active and are more naive to commonly used measures. We believe this will have great positive impacts on data collection if researchers choose to utilize it. Another option for researchers is to use Prime Panels, which has workers who are more naive to commonly used measures due to the size of the platform and its primary use for marketing research surveys which typically have very different data collection goals and uses different tools than those used on MTurk.

References

Peer, E., Vosgerau, J., & Acquisti, A. (2014). Reputation as a sufficient condition for data quality on Amazon Mechanical Turk. Behavior research methods, 46(4), 1023-1031.

Friday, December 1, 2017

Are MTurk workers who they say they are?

The internet has the reputation of being a place where people can hide in anonymity, and present as being very different people than who they actually are. Is this a problem on Mechanical Turk? Is the self-reported information provided by Mechanical Turk workers reliable? These are important questions which have been addressed with several different methods. Researchers have examined a) consistency of responding to the same questions over time and across studies b) the validity of responses, or the degree to which the items capture responses that represent the truth from participants. It turns out that there are certain situations in which MTurk workers are likely to lie, but they are who they say they are in almost all cases.


Consistency over time/Reliability:
One way of measuring truthfulness of responses is to examine how individual workers respond at different times to the same questions. In one study, data collected from over 80,000 MTurk workers examined the reliability in reported demographic information over time. This large study found that participants were overwhelmingly consistent when reporting demographic variables across different studies over time, with gender identification being 98.9% consistent, race 98.2% consistent, and birth year being 96.2% consistent, with this slightly lower score being largely due to technical issues rather than Turkers not being truthful (Rosenzweig, Robinson, & Litman, 2017).


Validity
Various forms of validity have been examined in data collected through MTurk, with results showing that data are by and large valid. We will focus on convergent validity, which refers to how a measure is correlated with other measures of known related constructs. Convergent validity of self-reported information at a group level can  be established by examining whether workers are providing logically consistent information. Data collected on TurkPrime show that associations between variables are consistent with what is found in the general population. For example, older Mechanical Turk workers tend to be more religious and more conservative, a pattern that is consistent with the general US population. The reported number of children correlates strongly with age and family status, as do divorce rates.  Self-reported time of day preference is correlated with the time of day that workers are actually active, which is also correlated with a cluster of clinical, personality, and behavioral variables that have been previously reported in the literature in studies of the general population (Unpublished Data). Similar consistent patterns have been observed in health information collected from Mechanical Turk workers. TurkPrime profiled over 10,000 Mechanical Turk workers on over 50 questions relating to physical health, with a factor analysis revealing that symptoms clustered around underlying conditions in the expected way. For example hypertension, high cholesterol, and diabetes formed a single factor. This factor, interpreted to be metabolic syndrome, correlated with other variables such as age and gender in the expected way. The rate of metabolic syndrome increases with age and was higher among men. BMI also correlated with self-reported exercise (See also Litman et al., 2015). Some other examples include the fact that rates of chronic illnesses are significantly higher among smokers compared to non-smokers, and strongly associated with BMI, with both higher and lower than average BMI being predictive of chronic illnesses.


Video tools that are currently in beta testing at TurkPrime are starting to be used to verify participants’ reported demographic characteristics such as gender, and race, and the presence of a second person for dyadic research, with promising initial results indicating that participants are highly truthful.


When Participants are Likely to Lie
Research has additionally examined the reliability of data collected when selection criteria were listed as a prerequisite to enter a study (e.g.“only open to males”). Data show that when participants are incentivized to not be truthful, such as when they are only able to take a lucrative study if they identify as a particular demographic group, they lie (Chandler & Paolacci, 2017; Rosenzweig et al., 2017). For example, in a study with a HIT title that said it was open “for men only”, 44% of participants who entered had previously consistently reported their gender as “female”.


Best Practices
When researchers want to selectively recruit participants on MTurk, they have several options. Some researchers recruit for a specific demographic group by including such specifications/selection criteria in a study open to all workers, and rely on workers to tell the truth when opting in to such a study. Based on the data, this is a mistake that will lead untruthful participants to opt-in. There are, however, several ways to selectively recruit participants that are who they say they are. One option is to use the qualifications system with characteristics already verified by MTurk, or by using TurkPrime’s qualification system. Another option is to run a study open to all workers, ask a series of initial demographic questions, and only have participants who match the desired demographic criteria proceed to the next round in the study, paying even those who were of the wrong demographic for their time. You can create worker groups on TurkPrime based on such pre-screenings which can help you track and subsequently recruit participants who match your criteria of interest.


References:
Rosenzweig, C., Robinson, J., & Litman, L. (2017, January). Are They Who They Say They Are?: Reliability and Validity of Web-Based Participants’ Self-Reported Demographic Information. Poster presented at the 18th Society for Personality and Social Psychology Annual Convention, San Antonio, TX.

Monday, November 20, 2017

Strengths and Limitations of Mechanical Turk


Hundreds of academic papers are published each year using data collected through Mechanical Turk. Researchers have gravitated to Mechanical Turk primarily because it provides high quality data quickly and affordably. However, Mechanical Turk has strengths and weaknesses as a platform for data collection. While Mechanical Turk has revolutionized data collection, it is by no means a perfect platform. Some of the major strengths and limitations of MTurk are summarized below.
Strengths
A source of quick and affordable data
Thousands of participants are looking for tasks on Mechanical Turk throughout the day, and can take your task with the click of a button. You can run a 10 minute survey with 100 participants for $1 each, and have all your data within the hour.
Data is reliable
Researchers have examined data quality on MTurk and have found that by and large, data are reliable, with participants performing on tasks in ways similar to more traditional samples. There is a useful reputation mechanism on MTurk, in which researchers can approve or reject the performance of workers on a given study. The reputation of each worker is based on the number of times their work was approved or rejected. Many researchers use a standard practice that relies on only using data from workers who have a 95% approval rating, thereby further ensuring high-quality data collection.
Participant pool is more representative compared to traditional subject pools
Traditional subject pools used in social science research are often samples that are convenient for researchers to obtain, such as undergraduates at a local university. Mechanical Turk has been shown to be more diverse, with participants who are closer to the U.S. population in terms of gender, age, race, education, and employment.
Limitations
There are two kinds of potential limitations on MTurk, technical limitations, and more fundamental limitations with the platform. Many of the technical limitations of MTurk have been resolved through scripts written by researchers or platforms such as TurkPrime, which help researchers do things they were not previously able to do on MTurk including
  • Exclude participants from a study based on participation in a previous study
  • Conduct longitudinal research
  • Make sure larger studies do not stall out after the first 500 to 1000 Workers
  • Communicate with many Workers at a time.
There are however several more fundamental limitations to data collection on MTurk:
Small population
There are about 100,000 Mechanical Turk workers who participate in academic studies each year. In any one month about 25,000 unique Mechanical Turk workers participate in online studies. These 25,000 workers participate in close to 600,000 monthly assignments. The more active workers complete hundreds of studies each month. The natural consequence of a small worker  population is that participants are continuously recycled across research labs. This creates a problem of ‘non-naivete’. Most participants on Mechanical Turk have been exposed to common experimental manipulations and this can affect their performance. Although the effects of this exposure have not been fully examined, recent research indicates that this may be impacting effect sizes of experimental manipulations, comprising data quality and the effectiveness of experimental manipulations.

Diversity

Although Mechanical Turk workers are significantly more diverse than the undergraduate subject pool, the Mechanical Turk population is significantly less diverse than the general US population. The population of MTurk workers is  significantly less politically diverse, more highly educated, younger, and less religious compared to the US population. This can complicate the way that data can be interpreted to be reliable on a population level.

Limited selective recruitment

Mechanical Turk has basic mechanisms to selectively recruit workers who have already been profiled. To accomplish this goal Mechanical Turk conducts  profiling HITs that are continuously available for workers.  However, Mechanical Turk is structured in such a way that it is much more difficult to recruit people based on characteristics that have not been profiled. For this reason while rudimentary selective recruitment mechanisms exist there are significant limitations on the ability to recruit specific segments of workers.


Solutions
TurkPrime offers researchers more specific selective recruitment opportunities, and has some features in development to help researchers target participants who are less active and therefore more naive to common experimental manipulations and survey measures. TurkPrime also offers access to PrimePanels, which has access to over 10 million participants, who can be selectively recruited, and are more diverse.


References:


Peer, E., Vosgerau, J., & Acquisti, A. (2014). Reputation as a sufficient condition for data quality on Amazon Mechanical Turk. Behavior research methods, 46(4), 1023-1031.

Friday, December 2, 2016

Dynamic Secret Completion Codes for SurveyMonkey

TurkPrime Supports Dynamic Completion Code for SurveyMonkey

Users of Qualtrics and Google Forms have long enjoyed the ability to integrate dynamic secret codes for each participant in their TurkPrime study. Dynamic codes ensure that each MTurk worker participating in your study receives a unique code completely eliminating the possibility that workers share secret codes.  In addition, with auto-approval enabled, those workers are automatically approved without any need for researchers to manually check the worker supplied secret codes and approve workers.

Now, TurkPrime users who host their studies on SurveyMonkey can also use dynamic completion codes as follows:
  1. Check off "Dynamic Completion Code For Qualtrics" when you design your TurkPrime study
  2. Redirect participants at the end of your SurveyMonkey study to https://www.turkprime.com/Router/DynamicCode

Friday, September 9, 2016

How to Create a Universal Exclude Worker List

Problem: 

Requesters may observe that some workers, even those with high Approval ratings, may not perform to their expectations on a study. Sometimes this may result in rejecting their work which affects the Worker approval rating. But, often the work is not acceptable for research but is not worthy of rejection, or, it may simply be the policy of the research lab to approve all assignments for IRB or some ethical standard they may follow. 

At this point the researcher may wish to exclude these workers from all future studies. MTurk has an option to Block a Worker (available through the API) but our experience has been that this solution is somewhat draconian and extreme: the effect of a Worker Block can trigger the suspension of the Worker's MTurk account. 

(Source: When I was a newbie Requester in 2012 I blocked some workers who gave me inconsistent and poor responses . I learned the hard way by having my Turkopticon rating suffer and the MTurk Worker discussion groups spread the bad word. I responded to the Worker complaints and Unblocked them to undo the damage to their reputation -- and mine!)

Solution:

Create a Universal Exclude List using the TurkPrime Worker Group feature. This exclude list simply excludes all specified workers in this group from taking a study with this Group Requirement. When you design your studies, just add this exclude group to your Worker Requirements and none of the workers in this exclude group will be qualified to take your study. 

This will achieve your goal of blocking undesired Workers without tarnishing their reputation.


Thursday, September 8, 2016

MTurk Panels on Your Own Requester Account

Studies with Panels  for just $0.15 - 0.75 / complete

Now you can run Mechanical Turk studies using your own Requester account and specify over two dozen demographic traits!.  The traits include gender, ethnicity, age, marital status and sexual orientation. But it does not stop there! The available options also include occupation, medical and health history, cell phone use and much more.


The cost ranges from $0.15 - $0.75 per completed assignment. For example, if you run a study with a panel of 100 White Males 40 and under with a cost of $0.42 / complete the TurkPrime Panel Fee is 0.42 * 100 = $42.00. The panel fee is determined by the incidence rate of your particular panel so that harder to reach demographics cost more...but are capped at a maximum of $0.75 per complete.

In addition, TurkPrime displays the feasibility of the study to run to completion. This is not a guarantee that the workers will take your study. since the MTurk workers, ultimately, decide whether they will accept and complete your study based on many factors including worker payment, requester rating on TurkOpticon and clarity of your study, among others. 

If you want to be certain your study will run to completion we recommend using the TurkPrime Lab Services of either Prime Panels or MTurk Panels where TurkPrime manages all user interaction, reaches out to workers to complete your study and guarantees the study will run to completion.

Studies with MTurk Panels will display the panel traits in the study dashboard along with a tag marking it as a panel study, as shown below.