This could just as well be framed as a taxonomy of magic or magical creatures or comic superpowers; you may disagree with details but the whole structure holds conceptually. There may be no actual facts involved (or maybe there are!), but the concepts are consistent. Also, I take this as a subset of the taxonomy of magic because there's (currently!) no scientific evidence but in the back of our heads we kind of feel like maybe we've experienced it or really really hope that there is some small ability there
First, let's define ESP (extrasensory perception) starting from examples, often being lucky enough that there are single English words that already capture the essence, and abstracting. There's clairvoyance (seeing the future), there's telepathy (perceiving someone's thoughts), mind-control (changing someone's thoughts by your own), speaking with the dead, telekinesis (moving objects with your mind), predicting random cards.
I'm setting an arbitrary boundary so that things we informally think are magical are not included (ghosts, gremlins, witches), that are 'obviously' unscientific and magic tricks (card tricks, optical illusions), which are intentionally supposed to seem magical but have a deterministic scientific explanation (astrology (depends supposedly directly on the location of the sun and planets)). These choices of mine are somewhat arbitrary. They could easily be included but then where do we stop (wait what about tarot and palm reading and tea leaves? what about entertainment magic, sleight of hand and actual tricks (ha ha that's hard to say right))
With these examples in mind, we can start to take apart what it means to be ESP and categorize all the kinds. The first thing to notice is that, along with perception, I am including action. So extrasensory perception or action is perceiving or doing things beyond our known senses. So we are well aware of seeing with our eyes and pushing with our hands; ESP is the ability to do those without currently known physiological organs. Presumably the organ will end up being the brain (the seat of thought), but maybe if we find out that we are able to see through the backs of cards using higher frequency receptors in our eyes (a deterministic scientific explanation) then this action will become a nonExtra Sensory Perception (NESP).
This brings up the tangent of making well formed categories. It is usually considered bad practice to have a subcategory, a sibling category, that is 'everything else that is not included'. For example, the category Vehicles could include Cars, Bikes, Planes, and NOS (Not Otherwise Specified). The latter category might cause difficulty because a sailboat will have to change category if a new subcategory of Vehicles, namely Boats, is created. (Note the difference between a category (eg Boats) and instances (sailboat), which of course could be generalized to become a category on its own)
A taxonomy of concepts forms a tree which expects all subtrees to be non-overlapping. Most collections of concepts end up having some overlaps, and this will be pointed out, but non-overlapping is a simplifying assumption that will make things easier to diagram.
- sensing
- 'perceiving' events
- clairvoyance, premonition - seeing events in the future, past, or remotely
- guessing cards
- predicting events
- telepathy - knowing others' thoughts
- mentalism - cold reading
- channeling - communicating with spirits
- seances - speaking with the dead (formerly actual people), knowing the thoughts of someone who has died
- sensing auras - 'seeing' the personality of a person
- out-of-body experience - astral projection
- acting
- telekinesis - or psychokinesis, moving objects
- levitation - raising objects
- oneself - as in extreme yoga
- somebody else
- objects
- making objects disappear
- modifying objects
- bending spoons
- destroying and remaking things (watches, dollar bills)
- pyrokinesis - starting fires (inspired/invented by fiction, Stephen King)
- telepathy - transfer of thoughts, more than just sensing
- sending thoughts, communicating
- putting ideas in someone's head
- mind control
- body control
I've never defined magic or science, only working with them informally. The creation of the relations among these things helps us define our terms, putting things together that go together but avoiding conflicts and inconsistencies by separating differences.
This is an exercise in philosophy and taxonomy. That is, I'm just playing with words and our mental perception of them, mainly because science could be done on these things, and has, but it has just never panned out. So all I have to go on them is what we imagine. So this taxonomy is not (as currently known) about scientific things, but is itself scientific because people have ideas of what these individual concepts could mean and could disagree with the relations I have put among them. Note that I've really only put a subset relation (is-a) and extremely minimal comments.
The only practical argument against any of these abilities being real (or scientific) is that no one has used any of these things for anything other than those particular entertainments. That is, if ESP/magic were repeatable with other objects, we could use, for example, the spoon bending skill for other metals and substances in industrial manufacture. Or we could teach quadriplegics how to do small tasks requiring dexterity. Or communicate without telephones. Of course the counterargument which is not a counterargument is the ability of pickpockets to take personal objects without us knowing. Some 'magic' is possible, just not by the purported skills.
What's interesting about the above taxonomy is that most (serious) people don't really believe that any of these phenomena are real. This is counting angels on a pinhead, building castles in the sky. There is no there there. But we've drawn a perfectly coherent picture. And frankly, it could turn out that some of these are physically realizable, through some sort of deterministic, scientific process.
Saturday, October 31, 2015
Friday, October 30, 2015
Language learners
What's happening with language learning:
- One unknown word ruins a sentence: a single unknown word in a sentence can totally negate any meaning the sentence might otherwise give. If you have never heard a word before, for a native speaker you often have enough context and history to figure the part of speech, how it relates to the other words, who is doing what to whom. But to the non-native learner, it totally throws off everything. All the other words, which you previously know, may now have their meanings in question. And all the intellectual energy you're expending trying to figure out the unknown word is taken away from all the other words, making the known words, still shaky in this new language, even less sure.
Native speaker reaction: "dumbledore must be a grocery store."
Learner reaction: "Did you just call me a ... a bird?"
- A learner is lenient, a native is strict: To a native speaker, there are lots of collections of words that are similar sounding and have similar meaning but are not the same. For example, all the various word forms of a conjugation "has, have, had" or cognates ", To the native speaker, these individual words are all very distinct. Using one instead of another is a glaring error, a discordant note, banging your thumb with a hammer obvious. To the language learner, they're kinda the same. To someone foreign to both, Italian and Spanish are a lot alike, you can sorta make half sense of both about the same. But of course to them they are mutually unintelligible (but can pick out a few words here and there). To the learner everything close is good enough. To the native the slightest hint of a difference is shockingly noticeable, strange, and almost unrecognizable. This works for all areas: pronunciation, syntax, word choice.
Native speaker reaction: "Want? That makes no sense. Did you want something at the grocery store? Did you want a grocery store? I don't get it."
Learner reaction: "Did you get chicken and potato salad?"
- Throw away step ladder: It seems universal in language teaching to start off with extremely simple sentences in the present indicative: "I read", "They eat". In English at least this is hardly ever used in practice. Surely there are short simple sentences that are actually used that can be taught.
What normal people actually say: "I used to go to the movies every Friday."
- Translationese: word for word translation is easy and often a sentence can preserve meaning from one language to the next with a constituent to constituent dictionary translation. But often "that's just not how they say it in X". Frankly in English that's just not they say it (see the throw away step ladder).
What is taught: "That is correct"
Natural: "Of course", "Right", "Yes", "Sure", "I guess so"
- Style is not grammar: lots of rules are given (a consistent single rule to learn is much easier to remember than a more complex one) for which it is actually a style rule or a rule of register (formal vs informal). Also, most language teaching is academic and for a future business or academic use. In most languages (moreso but still a little in English) there is a big difference between the language called X in school and that called X at home. This can lead to a good language learner to be 'better' or excessively more formal than a native speaker.
What people say: "I gotta go t'th'CVS 'n' pick up somethin' real quick"
Basic Pun "Where do polar bears vote?"
Learner: "The North Poll! Ha ha! I get it! Because 'poll' sounds just like 'pole' but they're two different things, one is for ..."
- Humor and sarcasm: anything other than the most literal will not be caught by the language learner. The learner probably has no idea that the same sounds can mean very different things depending on context (forgetting cultural background altogether). On the other hand, when a learner does notice a homophone, they will find it the most hilarious thing in the world but it will barely register with the native speaker.
Lack of humor "Where do polar bears vote?"
Native speaker: "The North Poll!"
Learner: "I don't understand. Can polar bears vote?"
Learner: "I don't understand. Can polar bears vote?"
Basic Pun "Where do polar bears vote?"
Learner: "The North Poll! Ha ha! I get it! Because 'poll' sounds just like 'pole' but they're two different things, one is for ..."
Native speaker: "Groan. Also, polar bears can't vote"
The point is that someone learning a language is using all their mental energy to pick out the right sequence of words, to get the order right, pronunciation, to remember that one weird word, etc etc that it's the most unnatural thing in the word, and they sometimes miss the eventual meaning.
Monday, October 26, 2015
Vapnik says "Deep Learning is the Devil"...maybe
Zach Lipton gave a summary of Vapnik's talk at Second Yandex School of Data Analysis conference (October 5-8, 2015, Berlin). Lipton wrote:
My interpretation of all this is that this is about the difference between science and engineering, or general vs specific. Coming up with a good general algorithm, I'm guessing Vapnik is thinking of SVMs or the idea of neural networks, is the study or science of ML, but most successes of Deep Learning (or really just particular and particularly large neural networks) come from the given design of the DL network.
As to clever vs brute force, somehow the statement that can be extracted is that DL is not clever but devilishly brute force. I'm not sure how to make sense of this (I don't see how DL is more brute force that SVM or logistic regression or random forests). Unless all the work that must be done in engineering a good DL is in creating the topology of nodes; this is not automatic at all but needs a lot of cleverness to make a successful learner. But the DL part enables that cleverness (which would otherwise be impossible).
Cleverness is not easily scalable; you can't just throw a whole bunch of extra nodes and arbitrary connections into a DL and hope it learns connections well, you have to organize the layers well. Those details,, the needed to be clever is what slows down the scaling and I am guessing it what is 'devilish' about DL.
This is all second hand and rewording of suggestions through someone's hearsay, and connecting dots that are barely mentioned and far apart. I'm totally putting words in his mouth, but this is what I expect Vapnik really means (or what I think Lipton thinks that Vapnik thinks, all telegraphically expressed). But really how much of anything is really not that?
Vapnik posited that ideas and intuitions come either from God or from the devil. The difference, he suggested is that God is clever, while the devil is not.and
Vapnik suggested that the devil appeared always in the form of brute force.and
[Vapnik] suggested that the study of machine learning is like trying to build a Stradivarius, while engineering solutions for practical problems was more like being a violinist
My interpretation of all this is that this is about the difference between science and engineering, or general vs specific. Coming up with a good general algorithm, I'm guessing Vapnik is thinking of SVMs or the idea of neural networks, is the study or science of ML, but most successes of Deep Learning (or really just particular and particularly large neural networks) come from the given design of the DL network.
As to clever vs brute force, somehow the statement that can be extracted is that DL is not clever but devilishly brute force. I'm not sure how to make sense of this (I don't see how DL is more brute force that SVM or logistic regression or random forests). Unless all the work that must be done in engineering a good DL is in creating the topology of nodes; this is not automatic at all but needs a lot of cleverness to make a successful learner. But the DL part enables that cleverness (which would otherwise be impossible).
Cleverness is not easily scalable; you can't just throw a whole bunch of extra nodes and arbitrary connections into a DL and hope it learns connections well, you have to organize the layers well. Those details,, the needed to be clever is what slows down the scaling and I am guessing it what is 'devilish' about DL.
This is all second hand and rewording of suggestions through someone's hearsay, and connecting dots that are barely mentioned and far apart. I'm totally putting words in his mouth, but this is what I expect Vapnik really means (or what I think Lipton thinks that Vapnik thinks, all telegraphically expressed). But really how much of anything is really not that?
Friday, October 23, 2015
Best Science Fiction Movies Ever
My list of best science movies ever:
I'm making little distinction between a single great movie in a series and the rest of the series. All of these refer mostly to the first ones in a series. I don't know what it is about Star Trek. The TV shows are way better than the movies; but the movies are somehow terrible (except as everyone agrees ST II: The Wrath of Khan). The modern J.J. Abrams reboots are enjoyable but nothing new and forgettable.
A handful require a mention but just don't make the list: THX-138; A Boy and his Dog; The Andromeda Strain; Stalker; A Clockwork Orange; Sleeper; Galaxy Quest; were all important when I was younger, but not really anymore (well, every other line in Sleeper is memorable).
Most scifi movies are just junk: action-adventure schlock.
The only one I feel bad about not putting on the list above is Wall-E. It's obviously well-done, there are a number of things in it that are memorable and prescient. But all I can say is "sorry animation".
Motivated by the Skeptics Guide to the Galaxy Episode #536 (10/17/2015) their top 5 sci fi movies.
- 2001: A Space Odyssey - the story is superior, the sets and effects still look modern.
- Blade Runner - again, the story is superior, the sets and effects still look modern. The attention to detail is amazing.
- Star Wars - simple minded but so audacious. Too many false notes to be the top
- The Matrix; Terminator - both stories excellent and very distinct but somehow similar. Hard to distinguish quality.
- Jurassic Park - almost too commercial
- Planet of the Apes - kitschy sets and costumes and dialog and acting (after the first one, they are all very amateurish), but rich in sci fi. This is the only one of mine that I think is very questionable. But I think it deserves a lot more credit.
- Road Warrior (and Mad Max I) -
- Brazil; Memento (I know they're not considered scifi, but both belong here for me)
- Too recent to judge reliably: Inception; Minority Report; Gravity; The Martian; District 9; Avatar; Children of Men; Source Code; Edge of Tomorrow; Looper; Predestination; Surrogates; Elysium; Limitless; Interstellar. I really liked all of these but I can't tell if they'll mean much to me later.
- 12/16 Arrival
I'm making little distinction between a single great movie in a series and the rest of the series. All of these refer mostly to the first ones in a series. I don't know what it is about Star Trek. The TV shows are way better than the movies; but the movies are somehow terrible (except as everyone agrees ST II: The Wrath of Khan). The modern J.J. Abrams reboots are enjoyable but nothing new and forgettable.
A handful require a mention but just don't make the list: THX-138; A Boy and his Dog; The Andromeda Strain; Stalker; A Clockwork Orange; Sleeper; Galaxy Quest; were all important when I was younger, but not really anymore (well, every other line in Sleeper is memorable).
I feel like I have to mention The Day the Earth Stood Still; Forbidden Planet; Metropolis only because they're near the top of everyone else's list. But I have to be honest and say they just feel so all around dated.
Most scifi movies are just junk: action-adventure schlock.
I went through a few top 100 lists just to make sure I wasn't leaving any out. If the title is not in this list (and there's lots), sorry, I just didn't think enough of it. If you were to ask me about one not here, I'd probably say, "I suppose it was OK, but it just doesn't fit in my best ever list". For example Tron. I remember being very excited about that as a kid, but, despite its main idea, it's just not that great. Some people think Dune (by David Lynch) was epic, but I feel like they must have seen a different movie (I thought the film was terrible) and their opinion was colored too much by the book. Soylent Green, Westworld, and The Omega Man - these are of the same kitschiness as Planet of the Apes, but only have one note each (also Charlton Heston mostly; is that the problem?).
The only one I feel bad about not putting on the list above is Wall-E. It's obviously well-done, there are a number of things in it that are memorable and prescient. But all I can say is "sorry animation".
At the end all I can say is that this is not opinion, it is objective truth. So if you don't agree, either you're wrong, or I have made a transcription error.
Motivated by the Skeptics Guide to the Galaxy Episode #536 (10/17/2015) their top 5 sci fi movies.
Tuesday, October 20, 2015
Cognitive computing is AI rebranded from the point of view of an app
Artificial Intelligence is whatever it is about computing that is sort of magic. As users we don't know why exactly it works but it just does. As builders it's like the dumbest of magic tricks: fast hands, misdirection, brute force beforehand. Sure a lot of research has gone into clever math for it, but once you look behind the curtain, you realize the excitement of the builder is in fooling the user, passing the Turing test by whatever means necessary. (I exaggerate considerably for effect. There's lots of rocket science behind the curtain, but what distinguishes it from algorithms is that it embraces inexact heuristics rather than shunning them).
AI doesn't have to be electronic. It could be mechanical, like a ball-bearing finite state machine that computes divisibility by three, or biological (like a chicken taught to do tictactoe). The chicken winning is magic. The brute force is determining the tictactoe decision tree then teaching the poor chicken. In the end, AI is mostly just computers.
It is usually something that humans are only able to do: language, vision, and logic.
AI is often used for just a heuristic such as the game 2048. The AI used to try to get a better score is essentially heuristics found by a good human player that were then coded as rules in a deterministic program. When you open the hood, there's no rocket science, it's just "always make a move that keeps the highest item in the corner", a human thought that is better than random and better than a beginner, but not perfect. An airline flight suggester is just (OK it is sorta rocket science) a special linear optimization problem (on a lot of data). Linear optimization is usually not considered AI but let's not quibble.
AI usually means that some learning was involved at some point but usually that learning is not continuous, learning in operation. Probably some ML algorithm was run on a lot of past data to generate a rule and then set in stone (until another pass on more recent data updates the rule).
The short history of AI is that it was invented/named in the late 50's, expected to solve all problems in a couple years in the mid 60's
Cognitive computing is IBM's way of reintroducing AI to consumers. It's not AI, but it's not not AI. That is, it is AI dressed up as usable applications or modules that can be fit together to give the appearance that a human is behind it without having an actual human having to step in to do it. The usual list of properties that a cognitive computing app has are: awareness of context of the user, giving the user what they want before they ask for it, they learn from experience, deals well with ambiguity. But then it will probably also incorporate human language input or visual pattern recognition or thinking through a number of inference steps.
Do you need an app that'll give you a new good tasting recipe for tacos? Deciding what's best is probably a good human task (but shhh we have an ML algorithm that figured out what are good ingredient combinations). Do you need an app that suggests to you good personalized travel plans? And now for something actually practical, do you need an app that will help discover cancer cures from buckets of EHR data?
All of these are not the usual single narrow one-off AI apps (back up a trailer, distinguish cats from dogs in images, compete in rock paper scissor competitions).
Without a doubt, cognitive computing is totally a hype/marketing term, new enough not to be some old over-used baggage-laden term like AI, not misleading (these are all sort of thinking apps) and vague enough to allow all sorts of companies to jump on the bandwagon with "Why yes, we've been doing cognitive computing before the term was invented!"
But hype terms can be useful. Cognitive computing is a good label for engineering a combination of features, some that are traditional AI and some that are just good design that are becoming more obvious to have. Whether it's the machines doing the thinking in silicon, or the engineers doing some extra thinking in design, as long as the machine looks like it's reading your mind then that's a good app. It'll be useful if it catches on.
AI doesn't have to be electronic. It could be mechanical, like a ball-bearing finite state machine that computes divisibility by three, or biological (like a chicken taught to do tictactoe). The chicken winning is magic. The brute force is determining the tictactoe decision tree then teaching the poor chicken. In the end, AI is mostly just computers.
It is usually something that humans are only able to do: language, vision, and logic.
AI is often used for just a heuristic such as the game 2048. The AI used to try to get a better score is essentially heuristics found by a good human player that were then coded as rules in a deterministic program. When you open the hood, there's no rocket science, it's just "always make a move that keeps the highest item in the corner", a human thought that is better than random and better than a beginner, but not perfect. An airline flight suggester is just (OK it is sorta rocket science) a special linear optimization problem (on a lot of data). Linear optimization is usually not considered AI but let's not quibble.
AI usually means that some learning was involved at some point but usually that learning is not continuous, learning in operation. Probably some ML algorithm was run on a lot of past data to generate a rule and then set in stone (until another pass on more recent data updates the rule).
The short history of AI is that it was invented/named in the late 50's, expected to solve all problems in a couple years in the mid 60's
Cognitive computing is IBM's way of reintroducing AI to consumers. It's not AI, but it's not not AI. That is, it is AI dressed up as usable applications or modules that can be fit together to give the appearance that a human is behind it without having an actual human having to step in to do it. The usual list of properties that a cognitive computing app has are: awareness of context of the user, giving the user what they want before they ask for it, they learn from experience, deals well with ambiguity. But then it will probably also incorporate human language input or visual pattern recognition or thinking through a number of inference steps.
Do you need an app that'll give you a new good tasting recipe for tacos? Deciding what's best is probably a good human task (but shhh we have an ML algorithm that figured out what are good ingredient combinations). Do you need an app that suggests to you good personalized travel plans? And now for something actually practical, do you need an app that will help discover cancer cures from buckets of EHR data?
All of these are not the usual single narrow one-off AI apps (back up a trailer, distinguish cats from dogs in images, compete in rock paper scissor competitions).
Without a doubt, cognitive computing is totally a hype/marketing term, new enough not to be some old over-used baggage-laden term like AI, not misleading (these are all sort of thinking apps) and vague enough to allow all sorts of companies to jump on the bandwagon with "Why yes, we've been doing cognitive computing before the term was invented!"
But hype terms can be useful. Cognitive computing is a good label for engineering a combination of features, some that are traditional AI and some that are just good design that are becoming more obvious to have. Whether it's the machines doing the thinking in silicon, or the engineers doing some extra thinking in design, as long as the machine looks like it's reading your mind then that's a good app. It'll be useful if it catches on.
Friday, October 16, 2015
The Long Burning Hype of Hype Indicators
(motivated by From Turing to Watson: The Long-Burning Hype of Machine Learning)
Hype is bullshit. It is not true, but it is also not necessarily false. In fact it has only a tenuous connection to the true/false dichotomy/continuum. It is only exclamation.
It is the hot-or-not score. It is the Time magazine weekly up or down cultural indicator. It is based on empty anecdotal perception, vaguely perceived frequency of mention or frequency of thought or coolness or I don't know what.
It is barely a measure of anything other than the .
The Gartner Hype Curve is also hype. It attempts to inform about the hype stage of many closely related items at once. But it turns out that is a piece of hype itself. The hype cycle is a well-hyped pseudo-scientific (non-evidenced based) proof-by-look-there's-a-picture.
Here is the general pattern:
It is very compelling. I have to be honest and say that that's exactly the timeline of how I think of things. At first I've just never heard of the thing. Then one mention.Then three in one day, then I hear and think of it all the time, then I get just sick of it, nauseated at the thought. Then it comes back as an accept everyday thing. Here's an example of a set of items from the
But... really? That's just a made up story. It seems to match what I think of as a story of popularity. It seems to match a good Hollywood drama: hero has early success and downfall and then third act of redemption.
And it is just the vaguist notion of mood swings. And also what is the point besides entertainment, or schadenfreude or rooting for a comeback?
Look at the following. So sciency. Look at all the data points that you can follow year after year (the data viz is a bit hard to read the course of any particular item, but that's a minor quibble in comparison to the the central problems).
The hype curve might be a useful thing, if only it measured something that is 1) coherent and 2) based on evidence. Introspection is a great inspiration but it is not measurable.
What is the meaning behind the graph? Also whatever the meaning what is the data underlying the graph?
The easy answer is the source of the data. It is simply the 'educated' guess of Gartner analysts. Not an actual number.
Take any particular item. Does it follow the curve? Does it match the icon? Each label is a hype term, but has its own definitional problems. Each term can be vague, have multiple meanings, and have multiple incommensurate sources.
Note also that the shape and timing of the graph is the same for all items. The different point icons are the only appeal to different scales for each item.
So time, the x-axis, is incoherent (unmentioned context based for every item). Different thing might move along the supposed curve at different rates.
But what about the y-axis? Is it popularity, that is, how often an item is mentioned (mentioned in tweets or on google)? or is it how successful an 'item' is (let's say quantitatively, money, earnings per year?) Even these ostensibly measurable concepts are problematic because of definition of terms (is one label the same or different than another).
And once you nail down what the measurement should be, Gartner isn't doing any kind of such measurement, and there's no guarantee that the hype curve shape is a common pattern. It may be that an item's curve is up then down (then dead). Or it may have a steady rise. Or it may have multiple hype peaks at different scales. Or frankly it may have a curve like any stock, up down, steady, with most any pattern imaginable.
If you want to make the hype curve useful. Pick a meaning (or meanings, heck go wild and have many types of hype) and then actually measure it. And only then will people... wel they won't accept that unconditionally, they'll also complain about the coherence of the concept measured and measuring difficulties. But at least it will be scientific evidenced based hype rather than just empty celebrity bullshit hype.
Hype is bullshit. It is not true, but it is also not necessarily false. In fact it has only a tenuous connection to the true/false dichotomy/continuum. It is only exclamation.
It is the hot-or-not score. It is the Time magazine weekly up or down cultural indicator. It is based on empty anecdotal perception, vaguely perceived frequency of mention or frequency of thought or coolness or I don't know what.
It is barely a measure of anything other than the .
The Gartner Hype Curve is also hype. It attempts to inform about the hype stage of many closely related items at once. But it turns out that is a piece of hype itself. The hype cycle is a well-hyped pseudo-scientific (non-evidenced based) proof-by-look-there's-a-picture.
Here is the general pattern:
It is very compelling. I have to be honest and say that that's exactly the timeline of how I think of things. At first I've just never heard of the thing. Then one mention.Then three in one day, then I hear and think of it all the time, then I get just sick of it, nauseated at the thought. Then it comes back as an accept everyday thing. Here's an example of a set of items from the
But... really? That's just a made up story. It seems to match what I think of as a story of popularity. It seems to match a good Hollywood drama: hero has early success and downfall and then third act of redemption.
And it is just the vaguist notion of mood swings. And also what is the point besides entertainment, or schadenfreude or rooting for a comeback?
Look at the following. So sciency. Look at all the data points that you can follow year after year (the data viz is a bit hard to read the course of any particular item, but that's a minor quibble in comparison to the the central problems).
The hype curve might be a useful thing, if only it measured something that is 1) coherent and 2) based on evidence. Introspection is a great inspiration but it is not measurable.
What is the meaning behind the graph? Also whatever the meaning what is the data underlying the graph?
The easy answer is the source of the data. It is simply the 'educated' guess of Gartner analysts. Not an actual number.
Take any particular item. Does it follow the curve? Does it match the icon? Each label is a hype term, but has its own definitional problems. Each term can be vague, have multiple meanings, and have multiple incommensurate sources.
Note also that the shape and timing of the graph is the same for all items. The different point icons are the only appeal to different scales for each item.
So time, the x-axis, is incoherent (unmentioned context based for every item). Different thing might move along the supposed curve at different rates.
But what about the y-axis? Is it popularity, that is, how often an item is mentioned (mentioned in tweets or on google)? or is it how successful an 'item' is (let's say quantitatively, money, earnings per year?) Even these ostensibly measurable concepts are problematic because of definition of terms (is one label the same or different than another).
And once you nail down what the measurement should be, Gartner isn't doing any kind of such measurement, and there's no guarantee that the hype curve shape is a common pattern. It may be that an item's curve is up then down (then dead). Or it may have a steady rise. Or it may have multiple hype peaks at different scales. Or frankly it may have a curve like any stock, up down, steady, with most any pattern imaginable.
If you want to make the hype curve useful. Pick a meaning (or meanings, heck go wild and have many types of hype) and then actually measure it. And only then will people... wel they won't accept that unconditionally, they'll also complain about the coherence of the concept measured and measuring difficulties. But at least it will be scientific evidenced based hype rather than just empty celebrity bullshit hype.
Friday, October 9, 2015
Kalman Filters = dynamic programming on linear systems for sensor accuracy
A Kalman filter is a method to increase the accuracy of a sensor in a linear system. (see this link for a visual explanation and derivation)
The usual example is for the position of a space ship. You have the position/velocity of the ship and also an independent sensor of those. Both of those are somewhat iffy (usually assumed for continuous variables to be Gaussian). Using these two iffy things together, you can get a much more accurate approximation (smaller variance than both) of the current position/velocity.
The other ingredient of the method that makes it get called a Kalman filter is the the change in the sensed data is expected to be linear, so that all of this can be modeled using simple repeated matrix operations.
Of course, this is not limited to dynamic mechanics but it makes the best presentation (because of the linear equations
The point here (which is not to explain Kalman filters) is that the computational method of correction (abstracting away the matrices) is one of a one step recurrence relation (new sensor data at each step too) which is essentially dynamic programming and even better, you only need to know the most recent item.
The usual example is for the position of a space ship. You have the position/velocity of the ship and also an independent sensor of those. Both of those are somewhat iffy (usually assumed for continuous variables to be Gaussian). Using these two iffy things together, you can get a much more accurate approximation (smaller variance than both) of the current position/velocity.
![]() |
| from bzarg |
The other ingredient of the method that makes it get called a Kalman filter is the the change in the sensed data is expected to be linear, so that all of this can be modeled using simple repeated matrix operations.
Of course, this is not limited to dynamic mechanics but it makes the best presentation (because of the linear equations
The point here (which is not to explain Kalman filters) is that the computational method of correction (abstracting away the matrices) is one of a one step recurrence relation (new sensor data at each step too) which is essentially dynamic programming and even better, you only need to know the most recent item.
What's the point of a hold-out set?
The purpose of a predictive model is to collect some sample data, calculate some function to help predict future unknown performance, hopefully with low error (or high accuracy).
The classic statistical procedure takes the sample, a small subset of past data called the data or for later purposes the training set, does some rocket science on that set (say, linear regression), produces the model (some coefficients, some small machine that says yes or no or outputs a guess on a single new data point), and maybe also produces some extra measures that says how good or bad the fit is expected to be (correlation coefficient, F-test). And we're done. So many papers and studies have been done over the years that follow this pattern.
But... what is a hold-out set? The modern way (not that modern) is to split the sample randomly into two parts, the training set (on which to do the classic part) and the test set or hold-out set to check. Run the model on all of the items in the test set and see how bad the fit. The test set is distinct from the training set because we want to validate on unseen data, we don't want to assume something we're trying to prove.
Why do this? It seems like such a waste. Why in a sense throw away perfectly good sample data on a test when you could use it in making a more accurate model? Why in a sense test again when you can use that test data to train? More data is better, right?
Well, you're not really throwing it away, but it does seem like a secondary, minor desire. After all, don't most statistical procedures compute some sort of quality measure on the entire set first? This desire not to 'waste' hard won sample data is very understandable; most of the labor in an experiment is not the statistics but in gathering the actual data.
Of course one could weakly justify this test set by saying it gives more reliable quality statistics.
The real desire for a holdout set is to combat overfitting. There are two sides to modeling: real life data is not perfect, the model is trying to get close to the rule behind the data, but it may go too far and get close to the data itself instead of the rule. The classic step gets us the first part, the modern step avoids going too far. A hint to the purpose is another name for the 'hold-out set which is validation set which name gives a better idea of its purpose. You create a model with the training set, and validate it with the validation set. You're validating your model, making sure that it does well what you claim does well. The first step in predictive modeling is to not underfit, to get close to reality that the data hopefully represents. The test or validation step is to make sure you don't overfit, get too close to the data at the expense of reality.
So I've weakly justified the desire for some kind of hold-out/test set. But how does one actually choose this set? Obviously a random subset but what size? The primary issue is a balance between the model and the goodness of fit: with smaller training set, more variance in the model; with smaller test set, more variance in the stats. There's no hard and fast rule (80/20 is considered reasonable).There are a number of strategies to deal with this.
A lot of this ignores the issue of what to do if your validation set has bad performance. What is the statistically 'right thing to do' then? Do you rejigger things knowingly? How adaptive can you be and avoid p-hacking? I'll save that for later.
The classic statistical procedure takes the sample, a small subset of past data called the data or for later purposes the training set, does some rocket science on that set (say, linear regression), produces the model (some coefficients, some small machine that says yes or no or outputs a guess on a single new data point), and maybe also produces some extra measures that says how good or bad the fit is expected to be (correlation coefficient, F-test). And we're done. So many papers and studies have been done over the years that follow this pattern.
But... what is a hold-out set? The modern way (not that modern) is to split the sample randomly into two parts, the training set (on which to do the classic part) and the test set or hold-out set to check. Run the model on all of the items in the test set and see how bad the fit. The test set is distinct from the training set because we want to validate on unseen data, we don't want to assume something we're trying to prove.
Why do this? It seems like such a waste. Why in a sense throw away perfectly good sample data on a test when you could use it in making a more accurate model? Why in a sense test again when you can use that test data to train? More data is better, right?
Well, you're not really throwing it away, but it does seem like a secondary, minor desire. After all, don't most statistical procedures compute some sort of quality measure on the entire set first? This desire not to 'waste' hard won sample data is very understandable; most of the labor in an experiment is not the statistics but in gathering the actual data.
Of course one could weakly justify this test set by saying it gives more reliable quality statistics.
The real desire for a holdout set is to combat overfitting. There are two sides to modeling: real life data is not perfect, the model is trying to get close to the rule behind the data, but it may go too far and get close to the data itself instead of the rule. The classic step gets us the first part, the modern step avoids going too far. A hint to the purpose is another name for the 'hold-out set which is validation set which name gives a better idea of its purpose. You create a model with the training set, and validate it with the validation set. You're validating your model, making sure that it does well what you claim does well. The first step in predictive modeling is to not underfit, to get close to reality that the data hopefully represents. The test or validation step is to make sure you don't overfit, get too close to the data at the expense of reality.
So I've weakly justified the desire for some kind of hold-out/test set. But how does one actually choose this set? Obviously a random subset but what size? The primary issue is a balance between the model and the goodness of fit: with smaller training set, more variance in the model; with smaller test set, more variance in the stats. There's no hard and fast rule (80/20 is considered reasonable).There are a number of strategies to deal with this.
- number not proportion - just make sure you have enough data points in each and after that proportion doesn't matter as much
- resample- do the test a few times on random subsamples. This is the very general procedure bootstrap/jackknife
- data partition and validate each as a test set against the rest - cross validation. This idea is to split the entire dataset into many samples and do the test/training on each set vs the rest. That is, all data is used as part of a training set and all as part of a test set at some point. There are many strategies here: leave one-out (LOOCV), where all but one is the training set and a single item is the test set, but do this for every single item in your data set. Under some models (like general linear regression models) you don't have to repeat the process n times because the math cancels out a lot (linearity is great!). Another method is k-fold CV where you split your data into k pieces (in practice often 5 or 10) and create a model on n-n/k items and validate on the, do that for each of these k pieces. It takes more time (k more times). LOOCV is essentially n-fold CV, so it is not efficient time wise when model creation takes a while (like for SVM)
A lot of this ignores the issue of what to do if your validation set has bad performance. What is the statistically 'right thing to do' then? Do you rejigger things knowingly? How adaptive can you be and avoid p-hacking? I'll save that for later.
Friday, October 2, 2015
Second- and third-order worrying
Once you're able to think consciously, once you're able to name things, you are then able to think about and name those same thoughts, think about thinking. Whether you're conscious of it or not. Consciousness is not necessarily self-awareness but it sure helps.
Some very basic behavioral phenomena could be considered higher order thinking. Altruism, doing something that is not in your immediate (or ever) personal best interest,
Let's say you have anxiety issues, panic attacks brought on by phobias, placed in trigger situations like a door being closed, meeting new people, or something as real life as big dogs. That is a first order anxiety (whether dysfunctional or very useful). A second order anxiety is if you're worried about losing your anxiolytic medication: you're afraid of being afraid.
Another situation: you're clinically depressed, and then you read an article that says that depressed people have a higher chance of contracting disease X (I made this up. I don't know this! Don't worry!). That could be considered grounds for being depressed.
Or paranoia... you could be playing rock paper scissors all day with imaginary adversaries!
So here is the problem. What does higher order worrying do for you?
The moral: Don't worry about third order worrying. After the second turtle (and it's turtles all the way down) they're all pretty much the same kind of turtle.
So that's one thing not to worry about.
Some very basic behavioral phenomena could be considered higher order thinking. Altruism, doing something that is not in your immediate (or ever) personal best interest,
Let's say you have anxiety issues, panic attacks brought on by phobias, placed in trigger situations like a door being closed, meeting new people, or something as real life as big dogs. That is a first order anxiety (whether dysfunctional or very useful). A second order anxiety is if you're worried about losing your anxiolytic medication: you're afraid of being afraid.
Another situation: you're clinically depressed, and then you read an article that says that depressed people have a higher chance of contracting disease X (I made this up. I don't know this! Don't worry!). That could be considered grounds for being depressed.
Or paranoia... you could be playing rock paper scissors all day with imaginary adversaries!
So here is the problem. What does higher order worrying do for you?
- Zero-th order: base activities. Let's say driving a car.
- First order: worrying about those activities. worrying about car wrecks, maintenance, getting lost. All these worries together is what anxiety is. Anxiety is not the next level, it just a set of worries.
- Second order: worrying about worrying too much. Worrying if you have an anxiety problem. I Am I worrying too much? Anxiety about anxiety.
- Third order: worrying about that. Here's the difficulty, what is 'that' and what does worrying about it really mean. Spelled out it is worrying about worrying about worrying. What it means is concerns about whether your concerns about anxiety are a problem. It's not that it is hard to think about (it is hard to think about), but that's not the point. The result though is that it is not an anxiety to have concerns about anxiety, that's not a thing, you just don't do it (also it is not a problem to have concerns over anyway). Your mind just doesn't go there (even after being led there). So it just isn't a concern.
The moral: Don't worry about third order worrying. After the second turtle (and it's turtles all the way down) they're all pretty much the same kind of turtle.
So that's one thing not to worry about.
Tuesday, September 29, 2015
Impenetrability of values of π
What is π?
Wait... before you spit out digits you memorized for that competition, what does it mean first? It is the ratio of the circumference of a circle to its diameter. That also seems a little weird because it's easy to measure a straight line against a straight line, but not against a weird curve. Rulers just don't work. But there's an easy way around that. Use a string, which is infinitely flexible, just make sure it is taut when measuring.
But it's still kind of weird. Who would've bothered to think to care about a ratio? (OK, it's easy to bother...you can walk around a circular castle, how long is it across?) But now that we've started thinking, π is a constant? That's also weird... why would you care? (well, it helps to simplify things and constants are simpler than variables, so...) It's obviously (once pointed out) more than 3 (inscribed regular hexagon) and less than 4 (circumscribed square). Anyway, it appears in quite a few mathematical places seemingly beyond geometry (eg the normal curve, the analytic continuation of the factorial: Γ(1/2) = sqrt(π)), for no apparent reason (well that's what inscrutability is all about!)
Sure, there are calculations with varying degrees of inscrutability: Archimedes' method (pictured), the Taylor expansions (Euler's π/4 = arctan 1, Machin's π/4 = 4 arctan 1/5 - arctan 1/239), the continued fraction expansion, Chudnovsky's method, the BBS spigot algorithm,
So I think I've dispensed with the impenetrability of π. I've made the tiniest of mathematical scratches and there's a long way to go. But I don't want to go there now. I want to explore the impenetrability of the value of π. The 3.1415926535... that value.
How big is π? Right, a little more than 3. But what is the point really? If you're shooting an arrow across the castle, 'a little more than 3' will do. If you're tying a rope across that 3.14 will do (and you'll want some slack anyway which will wash away any more digits.
Wait... before you spit out digits you memorized for that competition, what does it mean first? It is the ratio of the circumference of a circle to its diameter. That also seems a little weird because it's easy to measure a straight line against a straight line, but not against a weird curve. Rulers just don't work. But there's an easy way around that. Use a string, which is infinitely flexible, just make sure it is taut when measuring.
But it's still kind of weird. Who would've bothered to think to care about a ratio? (OK, it's easy to bother...you can walk around a circular castle, how long is it across?) But now that we've started thinking, π is a constant? That's also weird... why would you care? (well, it helps to simplify things and constants are simpler than variables, so...) It's obviously (once pointed out) more than 3 (inscribed regular hexagon) and less than 4 (circumscribed square). Anyway, it appears in quite a few mathematical places seemingly beyond geometry (eg the normal curve, the analytic continuation of the factorial: Γ(1/2) = sqrt(π)), for no apparent reason (well that's what inscrutability is all about!)
So I think I've dispensed with the impenetrability of π. I've made the tiniest of mathematical scratches and there's a long way to go. But I don't want to go there now. I want to explore the impenetrability of the value of π. The 3.1415926535... that value.
How big is π? Right, a little more than 3. But what is the point really? If you're shooting an arrow across the castle, 'a little more than 3' will do. If you're tying a rope across that 3.14 will do (and you'll want some slack anyway which will wash away any more digits.
The fraction 22/7 is often given as an easy approximation (= 3.142854...). But what's the point? to remember that fraction (3 digits plus where the division goes, or two small numbers) is just as much as the decimal explnsion 3.14. And 22/7 is somewhat impenetrable itself (oh fourth grade nightmares of fractions!) 7 goes into 22 ... argh how many times ... oh 3 plus what's left over .. hmm.. one seventh) Just use 3.14 and cut out all that nonsense.
The next best continued fraction expansion is 355/113. Six digits (slightly mnemonic in repeats, but really somewhat random) It gets you 3.14159292... seven correct digits (yes I'm fudging the rounding). But then do you really want to remember that fraction and have to divide 355 by 113? Yechh. Even plain old division is way more impenetrable than just a list of digits (as long as the calculation of those digits is correct). Essentially, I'm saying that 355 divided by 113 is just as impenetrable (in your head) as 4 arctan 1 (and roughly the same calculator button presses).
If you're measuring the circumference of the observable universe to the precision of the radius of a hydrogen atom (you never know when that might matter), then you really need at most 39 digits (46.6 billion ly = 8.8×1026 m / 5.29×10−11m ... oh looks more like 38 digits is enough (26 + 11 + 1?).
So what do you really need? If your castle of one kilometer in circumference is measured around in (oh I didn't explain why measuring around to calculate diameter rather than the other way... presumably you're in a siege around the castle and you need to shoot arrows all the way across. It's a use case that comes up more often than you'd think)...measured around with a rope to a precision of within a meter (oh I love decimal) 3 digits, a millimeter 3 more, and already the stretchiness in a kilometer of rope or errors in laying measuring tape a hundred times is way beyond a millimeter.
If you want to explore the digit patterns in π, then, totally, go for the spigot algorithm (oh yeah if you don't mind them in hex).
If you want to explore the digit patterns in π, then, totally, go for the spigot algorithm (oh yeah if you don't mind them in hex).
If you're measuring the circumference of the observable universe to the precision of the radius of a hydrogen atom (you never know when that might matter), then you really need at most 39 digits (46.6 billion ly = 8.8×1026 m / 5.29×10−11m ... oh looks more like 38 digits is enough (26 + 11 + 1?).
But if you're shooting arrows (or painting a storage tank), just use 'a little over 3' and you're gonna add a slop factor anyway because there's lots of little engineering give and take you have to compensate for.
Well, whichever, measure twice, cut once.
Well, whichever, measure twice, cut once.
(What would I do if I were programming a Mars lander? Hell, yeah, double precision or more!)
Saturday, September 26, 2015
More comparisons between Statistics and ML
This is a continuation of a post I made about differences between statistics and ML.
I'm not intentionally trying to piss people off ("How dare you imply that we are not as good as those other guys") but I suppose some things might be provocative and arguable. All generalizations are false but a dog with three legs is still a dog ("Are you calling me a dog? How dare you!"). Isn't the point here really that stats and ML have quite a bit in common? Also, I use 'data' as a mass noun "the data is consistent with an increase in effect". Like 'water', I use it grammatically as singular. So there.
Knowledge doesn't come to us in a package; it is discovered piece by piece, following the path of least resistance, with no overarching systematic plan to fill out. Afterwards, the stories are made coherent and clean. and oversimplified for the textbooks. Also, different people in different academic cultures may explore the same things but with different basic tools. Some people call themselves X, some call themselves Y, they both do Z. But X and Y never communicate, not because they are competitors but because their motivations, their culture, the building they are housed in on campus, are so very different, they just aren't even aware of the other's existence.
Statistics started in the 1800's with government and economic numbers, but then also sociology (Quetelet), and then at the beginning of the 1900's with agronomy (Fisher) before it then exploded in every natural science (medicine, psychology, econometrics, etc). Though it started from applications, the mathematics behind it (I blame Pearson?) came from mathematical analysis (all those normal curves and beta distributions are special functions of analysis). Everyday statistics is making hypotheses, doing a t-test, p-values, most likelihood estimators, Gamma distributions. The point of statistics is to take a lot of data and say one or two small things about it (x is better than y).
ML (machine learning), very distinctly, came out of the cybernetics/AI community, a mix of electrical engineers and computer scientists each of which have their own subcultures but closer to each other than they are to statistics. The mathematics behind ML came out of numerical analysis and industrial engineering, decision trees and linear algebra, linear programming. Everyday ML is neural networks, SVMs. The point of ML is to engineer automatic methods to take lots pf data (like pixels in a picture or a sound pattern) and convert that to a label (what the picture is) or text sequence.
The cultural overlap is basic data munging, data visualization, and logistic regression.
I think the primary social difference (which leads to a few technical differences) is the following. Stats is much older and has tried to solve a few problems very very well. They try to take as little data as possible (because they were historically constrained computationally) and determine knowledge. A lot of statistical consulting is judging the study design, determining what can be known with what probability and what assumptions (like prior distributions) restrict what can be known with what reliability. ML is much newer; expects lots of computational power. It often overlooks lessons learned by stats.
But then stats is a bit held back by its insistence on blind rigor. ML is creating techniques that very successful without worrying about the foundations, about what a p-value is a probability of, or whether it is a probability at all.
Machine Learning is almost entirely methods for solving prediction problems. Instead of a human looking through a set of data and eye-balling what the pattern is, let the algorithm look at way more instances than is humanly possible to get the pattern. Most of the methods are ad hoc: neural networks, naive Bayes, SVM, decision trees, random forests. There are no principles. Sorry, there is not the depth of principles that statistics has, except when it borrows those principles.
Machine Learning does include some learning techniques (in the Active Learning area where real time data feeds supply and modify the model), but is primarily a relabeling of Pattern Recognition (which is a more accurate name, somewhat closer to the prediction methods of complex models, the pattern in general being a very specific kind of model).
ML is concentrated in the AI section of a CS department or sprinkled throughout the engineering departments (robotics in MechE, EE (they do everything!). Or in real life in lots of industries, speech recognition, text analytics, vision.
Of course, there are some individuals who probably consider themselves in both camps (Breiman, Tibshirani, Hastie. What about Vapnik)?
In statistics, there has been a great internal controversy between frequentism vs Bayesianism. Frequentism is for lack of a better way of saying it, the traditional p-value analysis. Bayesiansim avoids these somewhat with the added controversial notion of allowing an assumed prior distribution set by the experimenter.
Less controversial though is the tension between descriptive statistics (or data exploration) and hypothesis testing.
ML is mostly Bayesian by default since arely are assumptions made about the distribution (or any investigation whatsoever about the effects of the distribution) and MCMC (Monte Carlo Markov Chain). The biggest controversy is between rule based learning and stochastic learning. The success of neural networks in the mid 80's (and the success of 'Google' methods in the 2000's) has largely killed rule learning except for maybe decision tree learning and association rules.
I'm not intentionally trying to piss people off ("How dare you imply that we are not as good as those other guys") but I suppose some things might be provocative and arguable. All generalizations are false but a dog with three legs is still a dog ("Are you calling me a dog? How dare you!"). Isn't the point here really that stats and ML have quite a bit in common? Also, I use 'data' as a mass noun "the data is consistent with an increase in effect". Like 'water', I use it grammatically as singular. So there.
Knowledge doesn't come to us in a package; it is discovered piece by piece, following the path of least resistance, with no overarching systematic plan to fill out. Afterwards, the stories are made coherent and clean. and oversimplified for the textbooks. Also, different people in different academic cultures may explore the same things but with different basic tools. Some people call themselves X, some call themselves Y, they both do Z. But X and Y never communicate, not because they are competitors but because their motivations, their culture, the building they are housed in on campus, are so very different, they just aren't even aware of the other's existence.
Statistics started in the 1800's with government and economic numbers, but then also sociology (Quetelet), and then at the beginning of the 1900's with agronomy (Fisher) before it then exploded in every natural science (medicine, psychology, econometrics, etc). Though it started from applications, the mathematics behind it (I blame Pearson?) came from mathematical analysis (all those normal curves and beta distributions are special functions of analysis). Everyday statistics is making hypotheses, doing a t-test, p-values, most likelihood estimators, Gamma distributions. The point of statistics is to take a lot of data and say one or two small things about it (x is better than y).
ML (machine learning), very distinctly, came out of the cybernetics/AI community, a mix of electrical engineers and computer scientists each of which have their own subcultures but closer to each other than they are to statistics. The mathematics behind ML came out of numerical analysis and industrial engineering, decision trees and linear algebra, linear programming. Everyday ML is neural networks, SVMs. The point of ML is to engineer automatic methods to take lots pf data (like pixels in a picture or a sound pattern) and convert that to a label (what the picture is) or text sequence.
The cultural overlap is basic data munging, data visualization, and logistic regression.
I think the primary social difference (which leads to a few technical differences) is the following. Stats is much older and has tried to solve a few problems very very well. They try to take as little data as possible (because they were historically constrained computationally) and determine knowledge. A lot of statistical consulting is judging the study design, determining what can be known with what probability and what assumptions (like prior distributions) restrict what can be known with what reliability. ML is much newer; expects lots of computational power. It often overlooks lessons learned by stats.
But then stats is a bit held back by its insistence on blind rigor. ML is creating techniques that very successful without worrying about the foundations, about what a p-value is a probability of, or whether it is a probability at all.
What they actually do
Statistics is the science of analysis of data: mean and standard deviation (descriptives, what the data looks like), distributions (eg normal, Chi-squared, Gamma, Poisson), p-values, hypothesis testing, type I/II errors, t-tests and ANOVA, regression and general linear models. Its foundations are probability theory which is applied measure theory which is applied analysis (distributions turn out to be mostly special functions). Concerns: significance, p-value, confidence intervals, power analysis, correct interpretation of data and inferences. There are principlesMachine Learning is almost entirely methods for solving prediction problems. Instead of a human looking through a set of data and eye-balling what the pattern is, let the algorithm look at way more instances than is humanly possible to get the pattern. Most of the methods are ad hoc: neural networks, naive Bayes, SVM, decision trees, random forests. There are no principles. Sorry, there is not the depth of principles that statistics has, except when it borrows those principles.
Misnomers
Both labels are misnomers. Statistics sure is used to study states and governments, but is overwhelmingly the province of (a very weird subset of) mathematics.
Machine Learning does include some learning techniques (in the Active Learning area where real time data feeds supply and modify the model), but is primarily a relabeling of Pattern Recognition (which is a more accurate name, somewhat closer to the prediction methods of complex models, the pattern in general being a very specific kind of model).
View from the outside
From the outside, statisticians are consultants for the research community for agronomy, econometrics, medicine, psychology, any academic science or applied version that takes a lot of data and (interestingly it is the softer sciences like psychology and sociology that send their grad students to the statistics departments for instruction, but the physicists and chemists, even though they may individually use some regressions, don’t usually depend on a statistician even thigh they may do a regression or two. Maybe they think they know enough to do it themselves?). Either way, ML people make more money, I don't know why.In industry (applied)
Statisticians are employed for quality control. This is their primary act as working statisticians. Taking samples of products, calculating error rate. ML people are more directly part of creating machines that do things in a fancy way, building things that work, like an assembly line robot for cars or zip code reader for handwritten mail.In academia
Statistics is concentrated in an academic statistics department (Often attached to a mathematics department or ag school) or as a group of consultants for agronomists or medical research.ML is concentrated in the AI section of a CS department or sprinkled throughout the engineering departments (robotics in MechE, EE (they do everything!). Or in real life in lots of industries, speech recognition, text analytics, vision.
Of course, there are some individuals who probably consider themselves in both camps (Breiman, Tibshirani, Hastie. What about Vapnik)?
Controversies
This has mostly been controversail as to what the differences are because of the tension between trying to assume they are the same but showing where the cultures make them different. Instead this is about the controversies within each.In statistics, there has been a great internal controversy between frequentism vs Bayesianism. Frequentism is for lack of a better way of saying it, the traditional p-value analysis. Bayesiansim avoids these somewhat with the added controversial notion of allowing an assumed prior distribution set by the experimenter.
Less controversial though is the tension between descriptive statistics (or data exploration) and hypothesis testing.
ML is mostly Bayesian by default since arely are assumptions made about the distribution (or any investigation whatsoever about the effects of the distribution) and MCMC (Monte Carlo Markov Chain). The biggest controversy is between rule based learning and stochastic learning. The success of neural networks in the mid 80's (and the success of 'Google' methods in the 2000's) has largely killed rule learning except for maybe decision tree learning and association rules.
Notation
Usually stats is the old fogie and ML is the uncultured upstart, but in mathematical notation it is different. ML, coming out of engineering, uses more traditional mathematical notation. Though nominally more closely connected to mathematics practice, statistics uses a bizarre overloading of notation that no one else in math uses. For probabilities, distributions, vectors and matrices. Every thing element has multiple meanings, context barely tells you what's the right syntax.Random notes
- ML is almost entirely about prediction, in stats there’s quite a bit else other than that.
- ML is almost entirely Bayesian (implicitly). Explicit Bayesianism is out of Stats. Frequentism, traditional statistics, is what most applied statistics uses.
- Stats is split into descriptive and inferential meaning either simplify the entirety of some data into a few representative numbers, or judge if some statement is true.. Descriptive creates patterns/hypotheses that then the inferential judges how good the patterns/hypotheses are
- predictions vs comparisons. ML is almost entirely predictive. Stats spends a lot of time on comparisons (is one set different from another, is the mean (central tendency) of one set significantly different from that of another)
- Leo Breiman also explained a distinction between algorithmic and data modeling which I think maps mostly to ML and stats respectively
How they're the same
I consider ML to be an intellectual subset of stats, taking a lot of data and getting a rule out of it no matter what the application. Whatever things get labeled ML, they really should have a statistical analysis (to be good), and statisticians should be willing to call these methods statistical. So what if they're in different departments.
Wednesday, September 23, 2015
Where is the universal electronic health record?
It's the 21st century. Where is our universal electronic health record? The one where all the medical knowledge about us individually is viewable by any doctor anywhere. You know, you get a yearly flu vaccine at your local drug store, and show up at the nearby emergency room for a sprained ankle, but when you go to your yearly checkup with your doc near work, they have no idea! Forget about it being possibly available when you're on vacation and get food poisoning and go to a non-local hospital.
In the middle of backest-woods China I can show up at an ATM for cash. On a flight 40,000 feet over the ocean I can get wifi to check on who was in that movie with that actress in that TV show. But in Boston, in the best place to get sick in the world, with every hospital connected with multiple medical schools, and every doctor with an MD and PhD and leader of the field that covers exactly your problem, you still have to, after getting a CT scan, walk down the hall to pick up a CD to physically deliver it yourself to your assigned specialist's office next door, nominally part of the same hospital network, but only financially connected, not electronically (oh, it is electronically connected, just not for that one thing. Oh, and the other things too which you'll have to walk back and get).
.
What's the point (other than that EHRs suck (and not just for the lack of interoperability))? The point is that the technology, the capability, and the knowledge to implement seamless connection for all electronic health data (images, reports, visits, medlists) was possible in the 70's ... with 60's technology. There is no rocket science here (a little electronics and programming sure). It is about as complex as ATMs. The internet should make things that much easier. But for whatever reason (oh there are reasons) it isn't there.
It is the year 2015, and there are plans to send people to Mars, so there is no technological reason why an interplanetary health record (IHR) doesn't already exist for use when they show up there. The record of the infection you got training in the desolate arctic landscape of Ellesmere Island. The dosimeter readings while stationed temporarily on the L2 jump-off station. Your monthly wellness-checkup with your PCP (well, remotely).
In the middle of backest-woods China I can show up at an ATM for cash. On a flight 40,000 feet over the ocean I can get wifi to check on who was in that movie with that actress in that TV show. But in Boston, in the best place to get sick in the world, with every hospital connected with multiple medical schools, and every doctor with an MD and PhD and leader of the field that covers exactly your problem, you still have to, after getting a CT scan, walk down the hall to pick up a CD to physically deliver it yourself to your assigned specialist's office next door, nominally part of the same hospital network, but only financially connected, not electronically (oh, it is electronically connected, just not for that one thing. Oh, and the other things too which you'll have to walk back and get)..
What's the point (other than that EHRs suck (and not just for the lack of interoperability))? The point is that the technology, the capability, and the knowledge to implement seamless connection for all electronic health data (images, reports, visits, medlists) was possible in the 70's ... with 60's technology. There is no rocket science here (a little electronics and programming sure). It is about as complex as ATMs. The internet should make things that much easier. But for whatever reason (oh there are reasons) it isn't there.
(that's not Jimmy Carter, it's a made up person for HIPAA compliance)
![]() |
| http://www.theplaidzebra.com/first-manned-mission-to-mars/ |
Right now all you get is your intraoffice electronic health record (that is, within an office, not between). It would work great if your PCP, endocrinologist, and cardiologist all belong to the same practice. Of course they don't. Sometimes you're lucky and a big hospital will be the only center for an area and all docs belong somehow to that one hospital. I'm not saying things are bad everywhere.
Wait. Expletive. I can't go to any local drugstore (again!) to get an over the counter bottle of Sudafed, some batteries for a game controller, and a jug of bleach for my socks without stormtroopers crashing through the windows, hog-tying me, and interrogating me on suspicion for running a meth lab (I mean every time), because I went to another drugstore across town for that very suspicious flu shot. At least somebody can connect systems. I was almost happy that they cared! About me!
Enough idle complaining. My idle blaming is that it is the health care businesses's fault. The docs are doing their job as well as they can. The businesses don't get anything out of making things easier on the patients or docs. I have all sorts of constructive suggestions just no one likes advice.
Friday, September 18, 2015
Confidence in association rules is identical to conditional probability
There's something that has bothered me for a while. In presentations of association rule learning (as a method of an unstructured learning method/data mining), the basic principles are:
And the various algorithms (brute force, apriori, Eclat, FP-growth) work on the list of transactions to discover association rules with high confidence. Confidence is the primary concept to be optimized
So what is the difficulty? That last definition of confidence. all that buildup with all that new vocabulary, all so straightforward and sensible, but all so new. There's something about ... confidence... that seems so familiar, but the notation... of implies and support .. it's just...
Of course this has been done elsewhere already.
Confidence is simply the conditional probability of Y given X. That's it. In notation:
which is the probability of Y occurring when restricted to when X is already known to have occurred (not temporally). What might be misleading here is 'and' versus 'union'. In the confidence formula we want the frequency of the itemset and in Pr we want the proportion of events. There is a just a little step of manipulating subsets and events here; the elements of the set unioned with those of Y is equivlanet to the event of those elements conjoined (= anded) with those of Y. A subset of elements S of T is the dual of the events T a subset of S.
Just a little rejiggering of notation and a whole set of concepts opens up to help think about the space of association rules.
(from Pier Luca Lanzi, DMTM 2015 - 05 Association Rules)
- the store - the set of all possible items {milk, bread, eggs, beer, diapers} = d
- transactions - a list of subsets from all possible items (a transactoin = 1 market basket) eg {milk, bread, eggs, beer}, could be represented by a 0-1 vector of length d. # of transactions = n <= 2^d
- itemsets - a subset of items in a transaction eg {milk, bread, eggs} or {bread, beer}, k-itemset has k items.
- support - support count = frequency of occurrence of n item set \sigma({bread}) = 2, support = proportion of an itemset to total transactions s({bread}) = \sigma({bread})/n = 1
- frequent itemset - itemset with s >= given threshold
- association rule: X-> Y,X,Y itemsets, The intention is that X implies Y, or if X appears in a transaction, Y is likely to appear also.
- support (X->Y) = \sigma(X \cup Y), fraction of transactions including both X and Y
- confidence - c(X->Y) = \sigma(X\cup Y)/\sigma(X), how often Y appears in transactions that have X
And the various algorithms (brute force, apriori, Eclat, FP-growth) work on the list of transactions to discover association rules with high confidence. Confidence is the primary concept to be optimized
So what is the difficulty? That last definition of confidence. all that buildup with all that new vocabulary, all so straightforward and sensible, but all so new. There's something about ... confidence... that seems so familiar, but the notation... of implies and support .. it's just...
Of course this has been done elsewhere already.
Confidence is simply the conditional probability of Y given X. That's it. In notation:
Pr(Y | X) = Pr( Y and X) / Pr( X )
which is the probability of Y occurring when restricted to when X is already known to have occurred (not temporally). What might be misleading here is 'and' versus 'union'. In the confidence formula we want the frequency of the itemset and in Pr we want the proportion of events. There is a just a little step of manipulating subsets and events here; the elements of the set unioned with those of Y is equivlanet to the event of those elements conjoined (= anded) with those of Y. A subset of elements S of T is the dual of the events T a subset of S.
Just a little rejiggering of notation and a whole set of concepts opens up to help think about the space of association rules.
(from Pier Luca Lanzi, DMTM 2015 - 05 Association Rules)
Robots having an Explosion, but not Cambrian
In my pursuit to eradicate bad analogies, the latest is in a paper "Is a Cambrian Explosion Coming for
Robotics?" by Gill A. Pratt in Journal of Economic Perspectives. It's a great paper, outlining reasons for an accelerating increase in the use of robots of all kinds and the technologies responsible for the acceleration, lots of enabling mechanisms (like energy storage improvements, combining learning in the cloud, wireless availability).
But to the metaphor. The Cambrian Explosion is first an explosion of varieties and then a very secondary implication an increase in incidences in the fossil records (lots more fossils). The usual explanation of the increase in fossils is that the newer life forms are more fossilizable, not that there are more individual lives.
Pratt's description of the explosion is not about varieties but about the technologies that will enable existing robots to be better.
I know this is a bit of a cavil because there were more fossils created during the Cambrian than before and could be called an explosion, but the usual provocative point about the Cambrian Explosion was that it was the great new variety that didn't exist before. Before the Cambrian, there were multicellular organisms (and fossil evidence of them), but during the Cambrian, lots of new anatomical structures seemed to appear for the first time (shells, tubes, etc).
The point is that when someone evokes 'Cambrian Evolution' it should be a metaphor for diversity not volume.
Otherwise, excellent article.
Tuesday, September 15, 2015
Driving in China
Just came back from a trip to China. Was driven around a lot, didn't drive myself. I noticed a few differences in driving style. In sum, I felt like I had to close my eyes a lot, which is apparently what the drivers do, too.
In the US, Canada, Europe, even France (!), people follow the rules of the road. They drive on one side of the street, they give pedestrians and cyclists a wide margin, even on dreaded traffic circles among the jostling there are rules of priority.
In China, the first impression is, as a backseat driver, to think 'Holy shit! Stop! you're going to hit that... whew... whoa you barely ran over.. whew... OH MY GOD you're going to kill us all... whew.. ' ad nauseam (literally). And the constant honking. I imagine I would be able to maneuver if only I could think but the incessant honking is so distracting.
In bigger cities, the wider roads have sectioned off parts of the road for bikes and scooters, presumably for safety. But whether these extra lanes are there or not, people on bikes, scooters, cars, trucks, etc will all intertwine.
In the US, there is the metarule, the rule of law, the rule that rules should be followed. Or if they're not followed a tinge of guilt and a speedy getaway. In China, the metarule is the rule of expediency, the rule that rules are there to guide you but really, I can fit right here at the moment, and look there's a pregnant woman on a scooter, with a young child sitting on her lap, and talking on a cellphone (she's not smoking that would be crazy), and she's making a left turn across the multi-lane intersection, yes, she can barely zip through before everyone fills the intersection, but oh, she cut across into the right turn lane of the crossing road and through three lanes of assorted vehicles, and left turn success! Some people are confident of what they're doing, some people not so, some people a little faster others a little slower, but everyone is aware of everyone else and they accommodate. Yes, I made some stuff up here. but just the cellphone. All the rest was faithful. Also, I was on a moped with two others (adults) in city traffic. But I'm here, without PTSD.
Sure they follow the traffic lights (as opposed to other countries where a stop light is very optional). If nobody is around sure they may slide through.
A slight detail that makes all this possible is that people just don't drive that fast. Not much faster than a moped (in traffic). That way everyone has enough time to make space for others and judgements about when to fill in that unoccupied space. On the highway however people will drive pretty fast, but there are hardly any cars on the super new clean highways.
There are cops everywhere, at every street corner in their cute little police boxes, but it seems they're not there for traffic but for shopping pedestrians. Also, the policemen seem more like bookish barely-out-of-college age accounting clerks, rather than the usual beefy, sunglassed, terminator-wannabes elsewhere.
In the US, Canada, Europe, even France (!), people follow the rules of the road. They drive on one side of the street, they give pedestrians and cyclists a wide margin, even on dreaded traffic circles among the jostling there are rules of priority.
In China, the first impression is, as a backseat driver, to think 'Holy shit! Stop! you're going to hit that... whew... whoa you barely ran over.. whew... OH MY GOD you're going to kill us all... whew.. ' ad nauseam (literally). And the constant honking. I imagine I would be able to maneuver if only I could think but the incessant honking is so distracting.
(from Living in China)
(that's really how it looks, but somehow that is not a traffic jam, just normal operating procedure, and cars get through)
But after a week of this, a pattern emerges. Not just the feel for the road, the different rules (and lack thereof), but also the different metarules. First, honking is not a mean thing. In the US, honking is like a rude gesture, an insult, a middle finger to your face. You do not use it unless 1) the light has changed and the person in front of you is an absent-minded enough idiot that they don't realize it's their turn or 2) Holy crap! My brakes are out and I'm coming towards an intersection or 3) Some mf- bastard just effing cut me off! In China, very much to the contrary, honking is a courtesy. Pardon me kind sir, I'm just a little behind you and I'm about to overtake you. I'm right here so be careful and don't swerve into me. Thank you so much!In the US, there is the metarule, the rule of law, the rule that rules should be followed. Or if they're not followed a tinge of guilt and a speedy getaway. In China, the metarule is the rule of expediency, the rule that rules are there to guide you but really, I can fit right here at the moment, and look there's a pregnant woman on a scooter, with a young child sitting on her lap, and talking on a cellphone (she's not smoking that would be crazy), and she's making a left turn across the multi-lane intersection, yes, she can barely zip through before everyone fills the intersection, but oh, she cut across into the right turn lane of the crossing road and through three lanes of assorted vehicles, and left turn success! Some people are confident of what they're doing, some people not so, some people a little faster others a little slower, but everyone is aware of everyone else and they accommodate. Yes, I made some stuff up here. but just the cellphone. All the rest was faithful. Also, I was on a moped with two others (adults) in city traffic. But I'm here, without PTSD.
(from Worst Intersections)
Sure they follow the traffic lights (as opposed to other countries where a stop light is very optional). If nobody is around sure they may slide through.
A slight detail that makes all this possible is that people just don't drive that fast. Not much faster than a moped (in traffic). That way everyone has enough time to make space for others and judgements about when to fill in that unoccupied space. On the highway however people will drive pretty fast, but there are hardly any cars on the super new clean highways.
There are cops everywhere, at every street corner in their cute little police boxes, but it seems they're not there for traffic but for shopping pedestrians. Also, the policemen seem more like bookish barely-out-of-college age accounting clerks, rather than the usual beefy, sunglassed, terminator-wannabes elsewhere.
Monday, September 14, 2015
Gödel not just good for incompleteness
This theorem is not provableKurt Gödel was famous for his incompleteness theorems (GIT) which entirely destroyed Hilbert's program (not really, just changed it's direction) and changed the face of philosophy of mathematics (probably should have but frankly not really), created recursive function theory and proof theory (pretty much).
But he is also well known within logic for many ground-breaking results there.
These results are
- proved the completeness theorem of predicate calculus (his PhD thesis, just before his incompleteness theorems, causing thousands of people-years in confusion because they refer to two different definitions of 'completeness, syntactic (for his PhD) and semantic (for GIT)
- created provability logic within modal logic (namely that Intuitionistic Propositional Logic is Interpretable in S4 )
- proved the consistency of the continuum hypothesis and the axiom of choice with set theory (just one half of independence of these two axioms from ZF set theory, Paul Cohen did the negative side).
I suppose there are other things that he did that would have made him famous if it weren't for each one of the above.
Sunday, September 13, 2015
First world problems, living in space, and calculators
Why do we exercise?
(Beware: this is a mix of opinions about space exploration, medical advice, and mathematical education, and social commentary, so pardon the whiplash.)
When I say 'we', I mean current first-world medical opinion is that daily exercise is important. Driving in cars, little walking, hyper-sugarfied drinks, large servings at restaurants, obesity cardiovascular disease, we are bombarded by personal advice and the fitness-industrial complex to exercise even if you have to drive to the gym.. Treatment for rich people diseases (and being in the first world nowadays allows some measure of curability/treatment either via surgery or lifestyle changes (diet and exercise)) are just not available in the third-world. They're just trying to get by, to make it through the day. They'd love to have the opportunity and control in their lives to eat more than they need or leisure time to rest, instead of having to walk 5 miles to get tainted water (or in inbetween countries, only get tainted water through the plumbing).
Many life threatening problems are so addressable by medical techniques that it is the relatively minor annoyances that have become sever medical crises in the first world, like Alzheimer's or social anxiety. The third-world is just trying to have subsistence level nourishment, not die from diarrhea or fever from infectious disease. (pardon my usage of the first vs third world terms. They are easier to distinguish whereas 'developed' and 'developing' are not).
Presumably before the industrial age (or becoming developed), people got lots of exercise walking around. Yet they died much younger. If it weren't for infectious disease, would they have had a longer life-expectancy?
It's not like exercise is some fancy new idea, it has been around forever. It's just that right now it is a public business. Even within the US, it is a bit of a rich vs poor distinction: those who have the money and time can exercise, but those who work two jobs and have kids really don't (one might ask if poor/busy people don't get lots of exercise naturally just by activity level...).
It has been well documented and studied that astronauts who spend lengthy times in space (weeks and months) as on the no-gravity ISS (International Space Station) have osteoporosis and muscle weakness. In order to counteract this they have as part of their schedule a rigorous exercise plan,
much more rigorous than on Earth. In fact, it has to be much more rigorous to account for the lack of gravity. The gravity on earth is naturally exercising us constantly. Just standing up on Earth we are using muscles from our legs and torso, even sitting is using your back. On the ISS, the zero-gravity environment is like lying down all the time. An astronauts schedule includes a couple hours a day on an exercise bike, or 'weight' lifting. (also, you can do a triathlon in space, but the swimming portion is hard)
You go to space and one particular feature of that environment which should be considered a great facilitator, the lack of gravity, has an effect on our biology which is expecting a much less lenient situation. And the biology pulls the other way. (I've heard that some-impact exercise like walking can be better than bike-riding, which is no-impact, because it encourages bone regrowth that counteracts osteoporosis. I've heard)
Which brings me to calculators in the math class (obviously). Or even computer aided algebra or automated proving systems for academics (and sometimes for engineering).
Kids these days, they can't even do long division! They've been coddled by calculators! How do we expect them to do science, let alone balance their checkbooks?
Calculators (and computer algebra systems to the nth degree) do the calculation for the user. Multiply two 10 digit numbers? A tedious exercise for a person, but a natural fit for a calculator. Solving that integral with square roots and trig symbolically? For a math/engineering whiz it's an hour long homework problem, but for the CAS, it's a natural fit. Solving it numerically? Insane for a person, but a natural fit again for the CAS.
The calculator (and CAS) is not intended to be a crutch that ends up weakening the user making them dependent. It makes you go that much further than you ever could go on foot.
A car takes us hundreds of miles in a day that we would never dream of doing on foot. If it makes us a bit lazy in taking the car for a few hundred yards, well, that's when we have to make sure we walk.
Technology puts us in the first world, but then we have to remember to exercise.
(Beware: this is a mix of opinions about space exploration, medical advice, and mathematical education, and social commentary, so pardon the whiplash.)
When I say 'we', I mean current first-world medical opinion is that daily exercise is important. Driving in cars, little walking, hyper-sugarfied drinks, large servings at restaurants, obesity cardiovascular disease, we are bombarded by personal advice and the fitness-industrial complex to exercise even if you have to drive to the gym.. Treatment for rich people diseases (and being in the first world nowadays allows some measure of curability/treatment either via surgery or lifestyle changes (diet and exercise)) are just not available in the third-world. They're just trying to get by, to make it through the day. They'd love to have the opportunity and control in their lives to eat more than they need or leisure time to rest, instead of having to walk 5 miles to get tainted water (or in inbetween countries, only get tainted water through the plumbing).
Many life threatening problems are so addressable by medical techniques that it is the relatively minor annoyances that have become sever medical crises in the first world, like Alzheimer's or social anxiety. The third-world is just trying to have subsistence level nourishment, not die from diarrhea or fever from infectious disease. (pardon my usage of the first vs third world terms. They are easier to distinguish whereas 'developed' and 'developing' are not).
Presumably before the industrial age (or becoming developed), people got lots of exercise walking around. Yet they died much younger. If it weren't for infectious disease, would they have had a longer life-expectancy?
It's not like exercise is some fancy new idea, it has been around forever. It's just that right now it is a public business. Even within the US, it is a bit of a rich vs poor distinction: those who have the money and time can exercise, but those who work two jobs and have kids really don't (one might ask if poor/busy people don't get lots of exercise naturally just by activity level...).
It has been well documented and studied that astronauts who spend lengthy times in space (weeks and months) as on the no-gravity ISS (International Space Station) have osteoporosis and muscle weakness. In order to counteract this they have as part of their schedule a rigorous exercise plan,
much more rigorous than on Earth. In fact, it has to be much more rigorous to account for the lack of gravity. The gravity on earth is naturally exercising us constantly. Just standing up on Earth we are using muscles from our legs and torso, even sitting is using your back. On the ISS, the zero-gravity environment is like lying down all the time. An astronauts schedule includes a couple hours a day on an exercise bike, or 'weight' lifting. (also, you can do a triathlon in space, but the swimming portion is hard)
You go to space and one particular feature of that environment which should be considered a great facilitator, the lack of gravity, has an effect on our biology which is expecting a much less lenient situation. And the biology pulls the other way. (I've heard that some-impact exercise like walking can be better than bike-riding, which is no-impact, because it encourages bone regrowth that counteracts osteoporosis. I've heard)
Which brings me to calculators in the math class (obviously). Or even computer aided algebra or automated proving systems for academics (and sometimes for engineering).
Kids these days, they can't even do long division! They've been coddled by calculators! How do we expect them to do science, let alone balance their checkbooks?
Calculators (and computer algebra systems to the nth degree) do the calculation for the user. Multiply two 10 digit numbers? A tedious exercise for a person, but a natural fit for a calculator. Solving that integral with square roots and trig symbolically? For a math/engineering whiz it's an hour long homework problem, but for the CAS, it's a natural fit. Solving it numerically? Insane for a person, but a natural fit again for the CAS.
The calculator (and CAS) is not intended to be a crutch that ends up weakening the user making them dependent. It makes you go that much further than you ever could go on foot.
A car takes us hundreds of miles in a day that we would never dream of doing on foot. If it makes us a bit lazy in taking the car for a few hundred yards, well, that's when we have to make sure we walk.
Technology puts us in the first world, but then we have to remember to exercise.
Thursday, September 10, 2015
What statisticians and ML'ers really think of each other
Labels aren't the thing, they just name the thing, and the same thing can have different names, and many different things have the same name. But people often take the label to be the thing.
'Statistics' and 'Machine Learning' are labels for two different things that have some overlap, not identical but cover a lot of the same things.
Statistics is concerned with averages and deviations, probability distributions, design of experiments, and regression, trying to extract knowledge out of tables of numerical data. The usual single sentence summaries are hardly distinguishable from many other things with data in their title, like databases or IT (Information Technology).
Machine Learning is a subset of Artificial Intelligence (itself considered a subset of Computer Science but practiced and motivated by other engineering departments and psychology related fields including linguistics, philosophy and neuroscience). It tries to extract patterns out of numerical data too, but has a different provenance. The two overlap some but each have their own separate culture and methods.
And more to the point, they’re really trying to do mostly the same things and the math for them both is often identical.
But what do they really think of each other?
From the point of view of the statisticians (people who call themselves with that label or are employed by institutions with that label) is that ML is a handful of ad hoc 'predictive analytics' done by a bunch of computer scientists, engineers, or amateurs (or worse!) pulling it out of their ass, their methods are immature (they don't know anything!) and don’t take into account the decades of principles established by the more mature staisticians for quality of results. That is, ML may do new, interesting things but they usually aren’t that new and they’ve never thought of all the methodological pitfalls that have been managed so well already by statistical principles (think of the data!). The statisticians may begrudgingly acknowledge that some of the ML methods are externally successful, but really, with such complicated models how do you know if it is any good outside of your toy domain when you haven’t done a proper analysis of your distributional assumptions? You ML people don't actually know anything!
People who say that they do ML probably do not give themselves the label statistician or work in a statistics group, but rather ‘are’ a computer scientist or engineer. Their point of view is that statisticians are studying pointless details about ancient brittle methods that aren’t particularly interesting, don’t really apply to all the new data sources, and just aren’t as good as this shiny new toy. Also, Bayes says p-values are dumb! The ML people may begrudgingly acknowledge that some of the statistical methods produce quality results, but really who cares about the normal curve and what about Bayes? You statisticians are so old and ossified!
From my point of view, it would be better for everybody if ML were considered a subset of statistics (but successfully studied in other departments) and ML methods could use a lot of analysis by statisticians. And a job that is labeled as data scientist should be easily fillable by a statistician or an ML person. Both sides need more exposure to the methods of the other.
See also Statistics and Machine Learning, Fight! (it's funding and conference culture) and Statistical Modeling the Two Cultures (by Breiman) (data vs algorthmic modeling), The Two Cultures: Statistics-vs Machine Learning for more opinions on the difference.
'Statistics' and 'Machine Learning' are labels for two different things that have some overlap, not identical but cover a lot of the same things.
Statistics is concerned with averages and deviations, probability distributions, design of experiments, and regression, trying to extract knowledge out of tables of numerical data. The usual single sentence summaries are hardly distinguishable from many other things with data in their title, like databases or IT (Information Technology).
Machine Learning is a subset of Artificial Intelligence (itself considered a subset of Computer Science but practiced and motivated by other engineering departments and psychology related fields including linguistics, philosophy and neuroscience). It tries to extract patterns out of numerical data too, but has a different provenance. The two overlap some but each have their own separate culture and methods.
And more to the point, they’re really trying to do mostly the same things and the math for them both is often identical.
From the point of view of the statisticians (people who call themselves with that label or are employed by institutions with that label) is that ML is a handful of ad hoc 'predictive analytics' done by a bunch of computer scientists, engineers, or amateurs (or worse!) pulling it out of their ass, their methods are immature (they don't know anything!) and don’t take into account the decades of principles established by the more mature staisticians for quality of results. That is, ML may do new, interesting things but they usually aren’t that new and they’ve never thought of all the methodological pitfalls that have been managed so well already by statistical principles (think of the data!). The statisticians may begrudgingly acknowledge that some of the ML methods are externally successful, but really, with such complicated models how do you know if it is any good outside of your toy domain when you haven’t done a proper analysis of your distributional assumptions? You ML people don't actually know anything!
People who say that they do ML probably do not give themselves the label statistician or work in a statistics group, but rather ‘are’ a computer scientist or engineer. Their point of view is that statisticians are studying pointless details about ancient brittle methods that aren’t particularly interesting, don’t really apply to all the new data sources, and just aren’t as good as this shiny new toy. Also, Bayes says p-values are dumb! The ML people may begrudgingly acknowledge that some of the statistical methods produce quality results, but really who cares about the normal curve and what about Bayes? You statisticians are so old and ossified!
From my point of view, it would be better for everybody if ML were considered a subset of statistics (but successfully studied in other departments) and ML methods could use a lot of analysis by statisticians. And a job that is labeled as data scientist should be easily fillable by a statistician or an ML person. Both sides need more exposure to the methods of the other.
See also Statistics and Machine Learning, Fight! (it's funding and conference culture) and Statistical Modeling the Two Cultures (by Breiman) (data vs algorthmic modeling), The Two Cultures: Statistics-vs Machine Learning for more opinions on the difference.
Subscribe to:
Posts (Atom)









