Showing posts with label business intelligence. Show all posts
Showing posts with label business intelligence. Show all posts

Sunday, February 14, 2010

Fraud detection and User Interaction: why are Millennials slower?

A scientist was conducting an experiment with a fly. He pulled off one of its legs and set it down to see if it could fly. Conclusion: a fly without one leg can still fly. He pared off a second leg and set it down, saying "Fly!" Conclusion: a fly without two legs can still fly. He removed all the legs and set the fly on the palm of his hand, shouting "Fly!" Conclusion: a fly without legs can still fly, briefly, before crashing to the floor. He pulled off all the fly's wings and set the fly on the palm of his hand, yelling "Fly!" Nothing. "Fly!" Nothing. Conclusion: a fly without wings is deaf.

This was an old, lousy and a bit vicious joke even when I was a kid. It does, however, effectively demonstrate a long lasting truth: it is not the collected data, but rather how we interpret it, that renders its effectiveness in decision making. Errors range from confusing cause and effect (is it that customers who experienced fraud are more active, on average, or that active customers are, in average, more prone to experience fraud?) to gross segmentation causing severe false positives; a lot of these cases are triggered by analysts sticking to high level, big numbers rather than complementing their analysis with case-by-case review and customer engagement. Business intelligence is a very important practice, and we must use our tools wisely to reach the best possible conclusions to guide our decisions.

One interesting case of interpretation I found was regarding Javelin's 2010 Identity Fraud Survey Report. Here's an excerpt from the link:

"18 to 24 Year Olds are Slowest to Detect Fraud – Millennials (consumers aged 18 to 24 years old) take nearly twice as many days to detect fraud, compared to other age groups, and thus are fraud victims for longer periods of time. Millennials were found to be the less likely to monitor accounts regularly and the least likely group to take advantage of monitoring programs offered by financial institutions. However, Millennials were the most likely group to take action such as switching primary banks or switching forms of payment."

Why is that? Well, looking for interesting opinions I came across this blog post. It suggests that Millennials are optimistic about the economy and feel invincible, being young, not imagining that fraud could happen to them. Interesting, but I don't buy into this kind of explanation, for two reasons: one, is that it's over simplistic in its description of Millennials' psych, but the second is that it puts a cap on our ability to engage with a group of users about their financials. It's just too important to let go: being able to engage with your user community to deter fraud will be a growing need for payment services in 2010 and beyond, and I claim that they expect this to happen. It just doesn't resonate with me that social networks and games can get you engaged but your bank or eWallet, the place where all your money is, can't. It's just a question of the right engagement model. What is the difference between those that work and those that fail? As a user myself, I don't feel like I have compelling interfaces that help me monitor my financials - and I log in to my online banking interface on a daily basis. There's just too much information, too many buttons and graphs to make sense. To add insult to injury, many monitoring programs (such as the lately advertized Chase debit card program) require users and parents to set their own monitoring rules. This reminds me of another area, online predator monitoring, which poses the same challenge to parents - you set the rules to monitor suspicious words in your child's IM. Seriously? We force the laymen to do our job for us? Can we really not provide a compelling, interactive, machine learning interface that provides an appealing user experience? I think we can. Especially if the alternative is accusing Millennials of being too optimistic.

Looping back to the beginning of the post, I'm just hypothesizing (or pulling the fly's leg, if you'd like). It's now a question of actually engaging with users and examining behavior to validate basic assumptions; something that we must do to make sure we understand the data we are getting. But this is my own hunch on Javelin's results. What do you think?

If you liked this post, please subscribe to my blog!

Monday, July 20, 2009

Ain't doing it right

"How many legs does a dog have if you call the tail a leg? Four; calling a tail a leg doesn't make it a leg." (ascribed to Abraham Lincoln)

In our business, to make a good decision, it is essential to know what really happned. So we discussed finding the single source of truth, but have not discussed ways for keeping it truthful. Oddly enough, the concept of immediate, detailed feedback is not as common as one would expect.

In your community of domain experts, the concept of "truth" should not only be determined but also enforced by members of the community. Note: not by a moderator; the members must know what the "truth" is (in procedures, in decisions and in deriving conclusions) but also be ready and empowered to call out their and others' mistakes. Because direct feedback is what enforces people to improve in the specific of their work. You do not only need people who can tell a tail from a leg - you need to give the one who detects it the means to show their finding to the general community.

This is not a matter of virtue, it's a matter of getting your business runnig the way it should. What happens if you under develop this area in your organization? Well, first you get only hindsight feedback, allowing you to know what's happening in delays of months and months (how much time does it take 90% of chargebacks to come in? exactly), but you also get feedback in aggregate levels (saying, for example, how many of person X's decisions were reversed) - meaning that you can't really find the trend and fix it.

I can't tell you it's fun - commenting, moderating or acting on the results of such feedback cycles - but one thing's for sure, it's way more effective than pretending your Risk experts live in DisneyLand. Giving and receiving proper feedback improves every bit of the cycle - and makes your business better at one of its core competencies.

Friday, June 12, 2009

Too much data, too little information

So, you have this big 1000 user system, with its flows and checkpoints and flags and pointers. If you've grown it well you have a dashboard showing you login numbers, counts of transactions, dollars moving around. You control it all from your NOC, pressing the little red buttons whenever necessary, moving dials and reading graphs. But the thing is, that seeing the bits and pieces of online life on your screen doesn't necessarily, and sometimes doesn't at all, help understand what's going on.

What IS going on in your system? What are users doing, and will that translate into the bottom
line? What can the numbers tell you?

Well, we've been through a few ideas. Experts knowledge ties symptomatic indicators with identities and with what they intend to do, so that you can at least start making sense. Collecting the data is one aspect, and using it to understand is a whole new area. When we reach tips and tricks on how to develop your own methodology, some of this might start ringing a bell. But this post is about one system that shouldn’t be adopted as your main tool if you’re the risk management expert – it’s about advising you to not count on hindsight based on business results.

No, no, don’t get me wrong – business results are important, one of the most important aspects of the business (and some will argue – the single most important – but that is another discussion). But using the bottom line (or even a highly detailed version of it, including a drill down of, for example, every auth rejection code) to indicate what the risks are in the system or worse yet – to indicate what needs to be fixed – is a call for bad judgment. Consider my favorite example, a hospital. If you needed to weigh two hospitals one against another, would you use the percentage of deceased patients as an indicator? Would it matter that one has an oncology department and the other doesn’t? Would it matter that one is in Mozambique and the other is in Mexico? Of course it would, since when all else is equal (in staff, training and tools – like your company compared to other retailers), fraud-on-entry (the hospitals’ location and the indigenous diseases you’d expect) and fraud MOs (the types of diseases that are actually seen and treated or not treated) have a big impact on the bottom line. Trying to use the numbers post risk controls, chargeback, CHB dispute and collections to understand what could have happened is trying to pin down a moving target – and the wrong one at that. Worse of all would be trying to design future systems based on the current snapshot, since you do not have any indication of what users do – just how much money it costs you, and user behavior is much more volatile than your incoming chargeback count.

When you come to understand what’s going on, business results are highly important. But letting them steer all of your team from looking at user behaviors will put you exactly where you don’t want to be – patching up holes in your system using a highly delayed hindsight mode. To be successful, combining data analysis and behavioral research is a must.

Tuesday, May 5, 2009

Differential diagnosis, people!

House - "Haven't done the MUGA."
Wilson - "Then how do you know she needs a heart transplant?"
House - "Got my aura read today. Said someone close to me had a broken heart."
(Season 1)

Yes, I admit it, I'm an avid "House, MD" fan. The fun part about this show is that a lot of people find meaning that's beyond the plain action to relate to - much different, I assume, than what the writers meant. Some watch it for plain medical aspect, like a good mystery story; some treat House as their fictitious mentor; some like the twists of the tale. I sometimes watch it like a tale of business intelligence and a general case of decision making with partial information.

Here's how it usually goes: in comes a case. It either looks suspicious upfront or bad indicators come up immediately at the beginning (by the way, did you notice that in most of the first half of season 1, it was seizures?). Then they go through "Differential diagnosis" and run various tests; additional symptoms are discovered, and usually the truth is discovered by connecting details that hid from the doctors (because "everybody lies") or simply because they didn't connect the dots.

Yeah, real life medicine isn't that simple, and sometimes even knowing what happened is too complicated to be nailed down case by case. Obviously catharsis doesn't come, like clockwork, every 35 minutes - just in time for the drama. But it's pretty similar, isn't it? In comes buyer A, and presents the details of person B. Not much to say about buyer A - their IP connection (anonymized?), their email (opened yesterday?), purchase details, maybe shipping address. Nothing much on person B either - name, address, credit card number. Would you let the purchase go through? Differential diagnosis, people! What test can we run to verify this person, or establish fraudulent behavior? What does it mean if they can verify the email, answer a call to their mobile phone, tell you that the issuing bank is Citi? What additional indicators are we missing? Because that's what the "game" is - in comes a case - what do you do? No one is dying, but your balance sheet is going to look pretty bad.

The trick about decision making in this case is understanding what the next step is. Our goal, whether asking the customer for additional details or looking for an additional data source (what's next - Family history review? MRI? CT scan?), is to reach a conclusion in as little steps as possible, meaning that we need to be able to choose the steps that contain as much information as possible. BI experts sometimes tend to get as much data as possible, sometimes at enormous costs (these external vendors don't come cheap). House's department costs the hospital millions of dollars a year, but that's human lives. We need to be cost effective.

One major way to work with this is automated decision making systems - expert system - which help experts reach decisions by dealing with the quantity of data by using statistical models for classification. Advanced systems, when correctly fed with symptoms (or fraud indicators), can even suggest tests to rule out corner cases. Constructing such a system is the end station of the long road that starts with the single source of truth - in House's case, the doctor. In fact, expert systems in the field of medicine usually outscore doctors in identifying illnesses based on differential diagnosis - it only makes sense, when you hear House's staff shooting diagnoses based on remarkable memory and years of experience. Which brings out the question - why doesn't House use one? It would immensely scale his ability to save lives.

But then again, how much fun will that be?

Tuesday, April 14, 2009

That one small detail

"When the Chinese government instituted the policy in 1979, it touched off a wave of sex-selective abortions as pregnant couples decided that if they could have only one child they would benefit most from having a boy. That helped leave modern China with the largest gender imbalance in the world. Today, there are 37 million more men than women in China, and many of the boys are growing up unable to find a job or start a family.

So what are these “surplus” boys doing to fill their time?"

This isn't just a story about risk management - it's a story of pure business intelligence - it is a story of freakonomics. The German police has spent years chasing down someone that turned out to be a phantom, a woman who wasn't really a feared killer in many different, distant crime scenes - but merely a lab worker whose DNA "slipped" onto the cotton swabs German CSI people used to collect evidence (on another note, wouldn't it be just morbidly funny if that person turned out to be a real-life German "Dexter" copycat?).

So what does an unsanitized cotton swab have to do with abortions in China, and with risk management?

When one approaches modeling of complex situations (either to explain what just happened, or to improve decision making in the future), often the "sense" made in the process gets deterred by the fact that not all the data is revealed. This is why when Freakonomics' author Steven D Levitt says something along the lines of "if we had enough data, we could unravel the mysteries of the universe", many of us nod (however, I must say, we are not always right); we are in constant search for the added detail that, when added to the equation, will help the story make sense. It's not only as extreme as claiming that a rise in abortions is correlated with a drop in crime rates - retailers are always looking for the additional factor that will verify a bank account, provide details for a phone number or do this automated super sophisticated AVS check. But fact is that most of the added data doesn't do the trick, since looking for that additional detail requires a system.

Yes, having a single source of truth helps give foundation, but even the brightest have a hard time without a system - and the right one at that - for collecting, validating and understanding data. I've seen this in organizations here and there and the German CSI story demostrates it well. The CSI department has a system for examining a crime scene and extracting evidence, and they came up with a concrete linking theory between cases. It didn't shed light on the actual identity of the misterious killer, however it gave an interesting spin to a bunch of unsolved crimes, until it didn't make sense anymore.

What the CSI department lacked was a key component of creating robust linking stories - indetifying common resources. That common BIN number in your last week's transactions might be a result of a data breach in the processor level, but might also be a result of a marketing campaign for a new eCard; and that repeated IP creating new accounts may be a script attacking your system but may also be a whole trend-struck fraternity house shopping through the same computer for that special item only you are offering for a great price. Noticing the trend, understanding it and making the right call on how to handle it are key decisions we are facing every day, and not only in eCommerce. Common resources are one simple example where correct classification, using an external resource, makes the difference between turning away good business and letting the fraudsters in; between chasing a phantom killer and tracking down a less-than-perfect lab worker. Using the right contructs for doing this is key in our ever-changing profession.