Thursday, December 31, 2015

What's with dictionary definitions for metaphorical usage?

Like a Jeopardy contestant giving an anecdote about their life changing experience when a pet dog tore up a favorite slipper, I have something I am terribly upset about.

I have noticed some online dictionaries giving metaphorical definitions. By this I mean that for a word, giving a meaning entry that is metaphorical, not its literal meaning.

For example, 'to devour'. Without checking, this means to eat ravenously. But it's easy to see that, say, a paper shredder could be said to devour some documents,

You may well note that many words, rather most words, really almost all words have multiple meanings (except for highly stipulated technical terms, and even then things can get loose). Our perception is usually that a word has one meaning and that's that. But then we notice that, well, that same spelling can be used for more than one distinct concept, usually nearby.

You may then well note that for many words, there really is a primary meaning: its meaning out of context that everyone thinks of first, and then secondary meanings, ones that appear in different contexts, that are slight extensions of the primary meaning, or used in analogous situations, not literally.

Here is the example for 'devour' from google:

de·vour
  • eat (food or prey) hungrily or quickly.
  • (of fire, disease, or other forces) consume (someone or something) destructively.
  • read (something) quickly and eagerly.
The first is the primary definition, the second a metaphorical one, the third... huh? That is definitely not what 'devour' means. Sure, one can easily use it in 'I devoured the sequel' meaning that I read the sequel quickly and eagerly. But that's not the meaning of 'devour'. That's not what 'mean' means. It's too specific. Does the omission mean you can't watch a movie voraciously? How come 'reading' is more devour-like than other metaphorical uses? This isn't right! If you include read, you should include every other possible metaphorical usage. But of course that is too laborious to imagine.

The difficulty I'm having is the demarcation line. When does a reasonable metaphorical usage of a word become dictionary-entry-worthy?

Taking the title word 'incensed', which was not deliberate, its primary and only definition is around 'angry', and no mention of the ostensible literal meaning which might have been 'burned like incense'. It already is a metaphor. The only definition is non-literal. So putting in metaphorical usages is necessary. At what point of semantic drift, at what point of leaving the original does a dying metaphor become dead, and at what point does the altered meaning move from quantitative difference to qualitatively requiring a new entry?

A close analogy is with suffixes. You can take any word in the dictionary and find some suffix that applies that will create a perfectly good word. 'Neologistically' is my favorite. 'Neologism' to 'neologistical' to 'neologistically'. Probably not in any dictionary, but perfectly understandable, sounds like a word, and is (arguably) undeniable as a word. Does it need to be in a dictionary? At what point do lexicographers decide not to include a possible variant?

There are a number of possibilities. Checking multiple dictionaries, most don't have the strange 'read' entry, only Google and Macmillan. What I suspect is that there is a tendency to require definitive alternate usage for an additional entry to be made when the entries are edited by humans. And that Google and/or Macmillan introduce metaphorical entries mechanically and its easier to be lenient. The latter two dictionaries certainly need human oversight; that is, the 'read' entry isn't a mistake but a lower threshold.

This will require looking into the editing policies of the various dictionaries.

Monday, December 28, 2015

Fix it now! Necessary minimal examples for change in English

The problems (two of them) are pronouncing can vs can't and fifteen vs fifty (or really more generally -n vs -n't and -ty vs -teen).

Well, it's not global warming, but there are a couple problems with English, specifically pronunciation, that are problems in information exchange. It's not some style thing, like the inarticulation of textspeak or emojis, or plain semantic errors like 'literally' for when it's not literal.

It's about pronunciation where the actual information being transferred is at it's most distinct. And the distinctions are being softened to almost disappearing but the information is as substantive as you can get. This screams out for language change, but...I just don't see any that people are unconsciously doing. The details are:

- can vs can't - In American English, standing alone, these two are identical except for the addition of the 't', which stands out in only the most labored of articulate speech. Word final 't' is unreleased in AmE, meaning that though 't' is usually aspirated elsewhere, but at the end of a word it is not.  This results in a difficulty for foreign language learners for hearing the distinction and for producing the distinction, what little there is.

But what's hard for foreigners can be hard for natives and even an area for language change (see the cot-caught merger).


  • I can do that
  • I can't do that


The intonation/timing for these is different (easy to see for native speakers but difficult for non-natives). But in high noise areas, these will be impossible to differentiate. "Can you move this 400 lb sofa off my chest?". "I can do that" vs "I can't do that". The answer is kind of important.

The same situation holds for a few other contractions involving 'not': could/couldn't.

My proposed solution: I can't think of a good natural one other than encouraging people to not contract when there is a 'not'. But for the positive version, the most natural way would by to shorten it considerably as '/kn/' or for both emphasize the sentence into nation "I cn do that" vs "I can't do that" (where the intonation would be terribly weird for the version of opposite polarity).

- -ty vs -teen:  The teens and the tens, thirteen/thirty and above, are similarly rooted but dictionary-pronounced distinctly. thir-TEEN vs THIR-tee. Different stress, different last syllable.

Except in regular speech these are about as close as you can get in pronunciation, and differently from most other close pronunciations (or eevn homophones), context rarely distinguishes them. "How far is it from the last to the next gas station?" "Oh about  thirt.. miles." "What? Is that thirty or thirteen miles?" "Thir-Teen" "What? THIR-ti or thir-TEEEE-nuh?" "The first one." "Oh. Then we're going to run out of gas."

The context doesn't say which one is more likely, and pronunciation has to be very articulate to tell the difference. The stress isn't pointed enough to tell. This is a perfect time for native speakers to unconsciously do something (make a slight change) to encourage better understanding (to avoid the waste of energy).

My suggestion? It's gonna sound bad but... for 'thirteen' say 'thir-TANE' or 'thir-TEEN-uh, and for 'thirty' say 'THIR-tuh'. To me these are the least unnatural and closest to natural pronunciation. Spelled out they look awful. But this is my prediction/suggestion.

I know. It sounds terrible.  But I don't want to run out of gas.


NB This is entirely about General American English... I can't expect to speak for other varieties of English. Also, I speak authoritatively about my own variety...but I may be biased and there may be free variation and some people may speak with more articulation than I can. So in one sense I am saying 'I am right' (which may be annoyingly authoritarian) but in another I am saying 'here is what I think I see, some others with great authority see it too, you can check now too now that you're aware of the possibility'.

Docs don't want to see your FitBit records, but ...

Docs don't want to see your FitBit records, but they do want your Holter monitor.

Docs do want you to exercise, or rather keep an active life and intense exercise is a good part of that but also make sure you don't sit down all the time, just walking and standing up and moving around calmly is probably a better thing to shoot for.

Having a FitBit or other exercise tracker, they consider that great, of course, because it means you are thinking about your fitness and trying to keep active.

But your PCP (your family doctor) does not want to see your FitBit records. They don't care how many steps you make a day. They don't care that you missed a few days last week (OK all last week. but the week before I made it every day!). They don't care about the fine detail of every second of every hour. They just want to know that you're using it. Or that you're making progress. Or that you're unable to hit the targets (knee surgery?).

Yes, the FitBit data is 'big data', and big data is useful...in general, but your PCP can't do anything with all those numbers. Your PCP isn't a big data consumer. Your PCP just wants to know (for the most part) Yes or No (or better than last year). Yes, totally get a FitBit, and totally because you think it'll please your PCP, and totally tell your PCP that you have one, and they'll totally be genuine proud of you. But don't think they can do anything with your data.

Maybe, just maybe, your cardiologist wants to know some of these details. Wait, no, your cardiologist doesn't want to know your FitBit data. What FitBit is tracking is not useful for them. Steps per day? Heart rate pattern over a time period? Oh...maybe a cardiologist does want to see a continuous EKG, like what you get from a Holter monitor. That is a specialist who can understand that data under a very prescribed situation. And with intensive study of that data as a specialist.

If say these fitness trackers somehow morph into a full body tricorder, measuring all possible metrics continuously, then maybe just maybe, it will become useful as a big continuous stream.

I'm sure the FitBit company data scientists could tell something interesting about behavior and fitness by working with sports physicians, epidemiologists, internal medicine in general, to discover interesting patterns of use of the product and specific health metrics.

But for you? Right now? Are you using a fitness tracker? Your doc does want to know that. Are you meeting your fitness goals everyday? Yes, your doc want to know that. Does your doc want to see your heart rate for the past month continuously by minute? No, your doc does not want to see that.

Wednesday, November 25, 2015

"Big ideas in mathematics"

My answers to the survey What do you consider to be the main "Big Ideas" in mathematics?

It's all about proofs:
- truth != proof (Goedel's Incompleteness theorems and what follow)
- logic and math still support each other (Reverse Mathematics)
- proofs are not just algebraic manipulation. They give meaning (Combinatorial proofs)
- automation in proof assistants will accelerate mathematical progress by humans (automated deduction)
- the most abstract of abstract nonsense (Category theory, no this not particularly about proofs)

These are all the biggest of big ideas in mathematics. All the others that are specific to an area are what I would consider too... contingent. Topology, algebraic geometry, algebraic K-theory, as important as they are, don't have any far reaching ideas. OK, number three above, combinatorial proofs, I just like a lot and is somewhat like those three I just disparaged.

Tuesday, November 17, 2015

Abused dataviz: periodic table and subway map

In addition to wordles, here are two more often abused data visualizations, the periodic table and the subway map.

Both diagram methods are intended to show that among a set of entities, there are many subsets, for the most part mutually exclusive but they have informative intersections. Think of the Venn diagram as the canonical diagram of subset relations. A subway map should have very few intersections (only a handful of entities are in more than one subset, the interchanges or transfer stations). The periodic table has a lot more structure, in fact, as a special case two-dimensional table, the full set can be split into mutually exclusive subsets in two distinct ways.

Take for example the original periodic table.

(from Science Notes)
What a great invention. Mendeleev compiled a bunch of disparate facts, similarities of elements, into a single visualization. The dataviz wasn't perfect, because there were gaps. But the picture was almost a theory, an extrapolation from data, that by 'testing' (further exploration) was confirmed by elements that fit nicely in those gaps. There have been attempts at organizing that chemical information in different ways but Mendeleev's holds primacy.

Nowadays, a periodic table is used for organizing a large set of items that have some similarities. Except the similarities have only tenuous systematic patterns. The point to the chemical table is that they fit nicely into rows and columns according to number of shells and number of electrons in outer shells (which predicts chemical properties nicely). The modern use of these periodic tables seem not to care what patterns in reality there are, just that pretty colors and list. Often the items in a column are not really related, and often they don't go from simple to weighty.


(from Expand via pinterest http://www.xpand.com.au/ )
In this example, the table is simply chart junk. The colors specify the mutually exclusive subsets but the rows and columns say absolutely nothing about the entities.


The periodic table of dataviz has some attempt at using the structure appropriately in the far left and far right columns, but in between its a mess. The site is great for examples, I'm only criticizing the use of the periodic table as the viz method. Note that almost all periodic table viz's use lockstep the funny unbalanced form of the table rather than fit it to the data (the properties of the entities). Instead the entities are shoehorned usually without any reason at all.

The point to a periodic table is that everything in a row should somehow be similar, and everything in a column should also somehow be similar. Also there should be some kind of progression from simple to complex down a column.

When is it appropriate to have a periodic table? When your set of items has two clear dimensions. There can be lots of gaps, or more in one position than another. But the two dimensions need to be clear. Also use those labels! Make sure everything in a column needs to be related. The rows don't necessarily have to be exactly related but at least of roughly the same complexity.
---




Subway maps are diagrams of connectivity of train systems. As a dataviz, they show that certain sets have a handful of points of intersection. Within a subset (shown by a line or track in the system) all the items are related. So when two lines intersect, that item must be a member of both subsets. Unfortunately, many 'subway' maps don't even bother with convention. They'll group items on a line that are only tenuously related, and then a 'transfer point' (an entity on two or more lines) ends up having little to do with either.

What makes a 'subway' map good is when the subsets have very few common entities. That will translate to only a few interchanges, making the diagram easier to create and less busy. It's a plus if you can order the entities along a line in a meaningful fashion (there is some inherent ordering).

Most uses of the 'subway map' dataviz, just like with the periodic table viz, either take a literal subway map (London's usually for obvious for obvious dataviz design homage) and shoehorn entities in, or make up their own but don't bother to make the lines and interchanges act like sets and intersections.



(from Becoming a data scientist) Using some domain knowledge, the items on each colored line aren't very coherent subsets, and their interchanges aren't really common between the two intersecting lines. There is quite a bit of overlap among these entities, lots of subset relations and intersections, but they are unfortunately not even bothered with.


A subway map viz is appropriate for a set of entities if those entities separate nicely into mutually exclusive subsets, with a handful of single entity intersections. If there are many intersections, then there are many constraints on how the lines meet each other.

Consider each entity as as having a list of features. If all the entities have a single feature that partitions the set (these are the subway lines) with very few entities with more than one line (the transfers) then the subway map is appropriate. If all the entities have two features, each partitioning the set in two distinct ways, then a periodic table is appropriate.

These dataviz strategies may well be meaningless chart junk simply to display a list with some structure. They are certainly esthetically pleasing (just like wordles!), but for the most part used irrelevantly. Most lists of entities are easily separated into sublists, with little extra structure, r quite a lot of complicated structure. The subway viz is good is there is a very little bit of common properties. The periodic table is good if there are two mostly coherent discrete dimensions, they don't have to be numbers.

The primary complaint is that the template is ostensibly knowledge based (scientific looking, 'sciency') but that the data poured into them just doesn't have that structure; the structure is a red herring. The dataviz should add something, should give you knowledge about the entities. If the items are on the same subway line, they should have some commonality. An entity in a periodic table should be similar somehow to the other entities in the same column and also the same row.

The alternative, when there is not enough structure, is a simple set of lists. If there is too much structure (lots of common features with little discernible pattern) is to use a venn diagram which captures all the possible intersections.

Or maybe I'm just complaining about incoherent sets and it's not even at the level of the top level structure being a red herring. It's a red herring that it's a red herring.

Sunday, November 15, 2015

In defense of publication bias and p-hacking


Publication bias and p-hacking have recently come under a few attacks even though it's not a new thing... not new ... at all.


Publication bias is the tendency to publish study results that have p-value (the statistical measure of 'significance') if the p-value is <= .05 (5%), the magical oversimplifying cutoff stated originally by Fisher as a.. well... magical oversimplification because he found that most people didn't really understand 'really' what p-values mean so he gave this as  a quick heuristic for significance.

'P-hacking' is the tendency for researchers to massage the data, the experimental design, the statistical method to improve the p-value output just over the threshold to 'significance' mostly to get around the publication bias.

These are problematic for different reasons. Publication bias leads the literature to ignore the non-information or weak information or non-results, things like 'X doesn't predict Y very well', or 'Treatment Z really doesn't really do much'.  These are things that would be nice to know, so a researcher in a field can then either avoid studying such unimportant things or discover refinements of the data that do show something important.

'P-hacking' is a bit more pernicious because it can lead in the direction of outright fabrication. Fixing poorly recorded data removing an outlier (somewhat reasonable, but it is controversial), but this is almost in the direction of fabricating data itself.

Given these well attested problems with publication bias and p-hacking, something about them is actually not so bad, in fact, they are an outcome of very reasonable and desirable scientific behaviors. This is not a justification for 'a road to hell is paved with good intentions', rather that some part of both is actually a good thing.

First, what is bad about publication bias (tending to publish only positive results)? From a literal, rational viewpoint, it is obviously denying half the story, creating all the false negatives. You should report all your results positive and negative to get a good picture. But that is a false equivalence. Positive results are not positive instances of a coin flip. Positive results are the interesting results. Interesting is something new and compelling. The null hypothesis is dull and lifeless. We already knew the null hypothesis. The null hypothesis is the air we walk through constantly. Reporting positive results is like pointing out a new pathway in the forest. Experts in the field, especially editors of academic journals, se many many results on slightly different phenomena. They have a good sense of what is new and important, and what has been done over and over again (and is maybe replication), but they also have a sense of what isn't important or isn't positive in the field. I'm not saying that negative results should not be published, but I do say that they don't need the boosting that positive ones do. A negative result is usually not that interesting. (In the natural sciences, that is; in more mathematical sciences, they are a different sort of thing, and often earth shattering)

And for p-hacking, sure, it is gaming the system. Hacking and gaming are things you do to improve something that are, let's say, hors du combat, outside the system Once you take a measurement, it is something to be gamed. Two runners competing for the fastest time? Train harder, eat better, lean at the tape, starting blocks, better shoes, bend the rules, shave your hair, make up new rules, take meds, get surgery, make up rules about these rules. The difficulty with experimentation and science is knowing what is cheating and what is allowable. For p-hacking, there are rules, things to do, things that are encouraged, things to avoid, and things you just can't do.

P-values are a third order measurement. First, data is the most primary measurement, a stopwatch, a rule, whatever. A statistic (like an average) is a measurement on data, you take a bunch of data and measure that set of data. For the mean, it gives you an idea of the center of the data. Then you can measure the p-value which is a measurement of how reliable the statistic is. At each stage of measurement, gaming can take place. You can manipulate data (remove outliers, 'fix' values) or manipulate the statistic (choose another, sample differently), manipulate the p-value (pick the best one, correct for multiple comparisons) and at each stage of gaming you can do it legitimately or not (no bias or much bias; yes, what the bias is is well underspecified).

The p-value itself is a time honored quantity associated with statistics. It is notoriously subtle and notoriously difficult to teach those subtleties. But it is a very useful measure of quality. It shouldn't be thrown away, just used carefully. When you mix dangerous chemicals, you do it under a fume hood, wear goggles, and have first aid nearby. When you calculate p-values, you make sure you don't calculate many on the same data and pick the best one.

Sure, sure, sure, publication bias and p-hacking are, as stated in that manner with their expected tendentious meanings, to be avoided as such. But the scientific process that results in those things, are not entirely evil. They are natural drives for knowledge and expression of knowledge and convincing people of knowledge. By saying 'natural' I'm not being lenient. Those drives have a correct part and an incorrect part. The part that we label bias and hacking are essentially bad, but the other part is not and is good. Not everything is bad about them.


Wednesday, November 11, 2015

Progress in making Star Trek tech real

There's been a lot of cool sci-fi technology over the years in Star Trek: transporters, phasers, faster-than-light travel. And by sci-fi I mean 'convenient but impossible stuff that helped get the plot move a long'. Star Trek TOS introduced a number of things, NG a few more, the movies I can't think of anything more than "an excess of chronotron particles has created an anomalous rift in the space-time continuum".


But it's kind of funny - since the original series, engineers have almost taken these sci-fi things as a challenge. "Star Trek can do it in the 23rd century. I will make it happen now!". Some things have actually happened; some have been found to be physically impossible. And all sorts between.
Here is an inventory of these plot devices/technologies and 'our' progress (star date spring 2015, 50 years after TOS). 
  • transporter - (from The Guardian). It (presumably) records all the positions and velocities of all particles in an abject, and recreates them ... elsewhere. This is theoretically possible and experimentally shown to work on individual electrons. But it needs a lot of work to scale this up to safely transfer humans. It just seems computationally infeasible and would take inordinate amounts of energy to make sure that all the atomic particles get transferred to just the right place. Some scenarios of this require that you have to kill your own doppelganger (it's really just a duplication device). Also, there are myriad niggling details like accounting for the relative speeds of the source and target locations (planets and satellites spinning around at huge velocities). However, one of the side benefits of transporter technology is that it allows you to remove infectious diseases and compute your genome, which, if the quantum relocation problems are solved, would surely be a piece of cake to do. But might allow mishaps aplenty as in many Startrek subplots, and other movies like The Fly. Not impossible, but very difficult with today's technology and economics. There are claims that transporting of objects has been done (see The Guardian article), but I consider that cheating of a form that just isn't cricket.That's a form of manufacturing (which may suffice for the entry on replicators).



  • automatic door opening - https://www.youtube.com/watch?v=6CSmkym-Stw  This is old news. Every grocery store has had these for years. Either through motion sensors, RFID chips, or weight pads
  • artificial gravity - this is a difficult one. What does it mean to have this technology? Without a large mass, can you simulate attraction to a flat surface? Or does a spinning cylinder which mimics gravity count (you'll be able to walk mostly normally, but water will still form messy blobs instead of pouring straight down)? The latter gets a lot, but the former...I don't think there is any physics that says this is at all possible. There's no graviton generator.

    However, one can cheat reasonably here. You can 'simulate' gravity on a surface as long as you make the surface accelerate towards the thing you want to exhibit gravity against. Make you spaceship accelerate towards a destination at 32 ft/sec/sec and It feels just like Earth! (turn around halfway and slow down at the same rate, because deceleration is just acceleration in the other direction. Note that this requires lots of fuel, fuel the whole way, rather than just coasting.

    Or the old fashioned spinning cylinder would work (with a large enough cylinder), going around a circle is acceleration, too.

    Neither of these is cheating, those are real equivalents of gravity. But if what you want is localized gravity, say over a square yard, as different from the adjacent square yard, then no, that's not going to happen.

    Progress score: by thinking of gravity as acceleration then yes in some contexts. But, arbitrarily, no, physically impossible.
  • invisible force fields or shields (for protection in battle) - same thing. This is actual magic. In that science is not involved. OK, this might be cheatable with lots of magnets. Progress score: never
  • cloaking device - The ability to make yourself invisible. Or maybe they've rethought what it means into a solvable formation. Yes, there is technological progress on this, wrapping light to go around an object, or camera on one side projecting to the other. Progress score: sometimes totally yes, other times in progress.
  • tractor beam - If there were arbitrary artificial gravity, it could be slightly modified to produce a tractor beam. Except I've heard of laser/magnet/ultrasound pincers for very small scale manipulation in air of particles (like dust). Progress score: not at all. Cheating: maybe
  • phasor (weapon for stun or kill) - sorta, not really. Like solar energy, there is the basic technology there, but it just hasn't progressed well. There are 'energy beams' but they take a lot of energy to produce, which is probably too difficult to engineer down to a handheld. But what about those green pointer lasers? I guess making you close your eyes is something. Progress score: sort of, very slow
  • photon torpedo - I don't know what's inside that thing. But we have had bad enough destructive bombs since before Star Trek came out. Progress score: already done, but maybe it doesn't glow or have flashing lights on it as it approaches its target, probably for good reason.
  • faster than light travel - This is obviously theoretically impossible, even for information, forget objects, All that wormhole stuff is nonsense (at least wormholes that allow an entire ship to travel through unscathed). Dude, if you go into a black hole, you'll be ripped apart before you even get to the event horizon. Of course there may be some cheating thing (transporters?) Progress score: never
  • computer disks -

    from filmjunk) quickly surpassed and obsoleted. I love this one as an example because it was one of the mini minor details in the TV shows that really made it great. Instead of big file folders of paper. "Just a small piece of plastic? Wow, the future is crazee!." That was the 60's. Then with PCs in the late 70's/early eighties, the floppy disk could be used as a replacement for paper. and it slowly got smaller until the early nineties the 3.5 in disk was widely used and pretty much identical in function to the ST disk. By the mid 2000's, we have thumb drives, barely noticeable on your keychain. And now people keep everything in the cloud. We don't even need a physical device to keep our files, it's just there in the ether. ST had the idea, people created it, people surpassed it, no one uses physical devices anymore.
  • communicator - Cellphones were invented in the '80s and have improved ever since, even better with apps (smart phones). I guess there are some limitations like we can't talk to the ISS from the ground with our phones (maybe someone at NASA can?) Wait.. we can tweet to the ISS on our phones. Progress score: Done.
  • talking computer - yes, to great extent. It's not perfect.It doesn't do all languages automatically. But for European languages, it gets syntax  and word choice mostly right (except for a few dings). Slow but steady progress by scientists over the years have created this (despite the bravado of the earliest years of AI claiming it could be done in a couple of years)
  • tricorder (handheld noninvasive medical analysis) - not yet, but getting there. There are all sorts of tools now for blood sugar, blood oxigen, temp, etc. Also, radiology is slowly miniaturizing. Next hurdle, hand held DNA sequencer?
  • universal translator - crazy impossible AI in the 60's. Hardwon but infinitesimally incremental progress over the years has recently resulted in a passable intermediate stage by Google. Kind of the same explanation as Siri. It's not the best but it's workable if you don't talk too fast.
  • DNA analysis - I mentioned the genome when I discussed the transporter and tricorder. Currently, if you have the locations of all the atoms in your body) presumably one can then go through some process on that data to get your DNA sequenced. But without (which the tech to produce the transporter would allow) genome can be done by a big machine for <$1K, but analysis and miniaturization not yet.
  • time travel - OK everybody, just shut up. Time travel is not possible at all. At least not in any way like in popular fiction. Wait, I'm sorry, Planet of the Apes got it perfectly right. You can travel into the future (we're already doing that right now, right?), but faster than normal. That is, if you travel at a non-trivial percentage of the speed of light, your time will be slower than others, and when you come back, everybody will have aged more than you. Other people will think you popped in from the past. That's about the extent that physics allows time travel. That's it. No going back in time (we'd have noticed people doing it already), and if you could you couldn't do anything that hasn't happened already. It's all just a plot device that appeals to successful primate brains. When the rodents take over (look at their hands!) and start telling stories, they'll come up with it too.
  • replicator - not really, but minimal progress. I consider 3D printing to be one part of this for which there has been great progress (and still more to go). The 3DP will create the physical substrate for an object (the parenchyma if you will). The missing part is the chemical or substance part. If you want to replicate a ham sandwich, you can (could, with some work) 3DP the shape and consistency of the bread, tomato, lettuce, ham, and mayo. But for each of those you still need the taste. And that will need some chemical engineering that I am not aware of. Subtleties in taste, not everything in the grocery store can be imitated with artificial flavoring. Progress score: early stages of prototyping
  • holodek - we're almost there. In the 90's there were rooms called 'The Cave' where your environment was displayed on the walls, and you held onto a control device that noted your orientation and movement. It was neat. On the way but not the best. The gaming community has created very life like images and environments. If you call it virtual reality there are VR helmets now. Not totally VR, not total immersion, but even closer. Basically what is needed is better plots and virtual acting.
Did I forget some glaring ones? Sure, the communicator. We've done it. Next.

And surely there are a lot of technologies mentioned before and after, not in ST, that are worthy of discussion. But this is about ST.

For many of these, where I say no or impossible, I'm pretty much challenging you to go around it, by rethinking the rules or outright cheating. That's how engineering works. To get from A to B better than a horse or bird is not to add more legs or implement wings. So maybe you can make a levitating car (but why?) as long as you have a track or ferromagnetic surface. 

Monday, November 2, 2015

Language to ask about language while learning

What if you had to learn a new language from scratch as an adult? (because kids just automagically pick it up) and you can't speak English because the other person doesn't know English.

You can always point a finger to get basic nouns, but at some stage (like lesson two) the nouns will get too abstract to point at. You can point at different examples of food but not 'food' itself. Likewise some verbs you can show, "I am running", but they'll get abstract even quicker "I have".
I've collected some phrases that should be a special section of its own in language learning. They are intended to help you bootstrap into a language, if you are so lucky to find yourself in this intermediate scenario.

You really need two kinds of things, a set of bootstrapping utterances that a normal fluent adult (works fine for kids too) speaking the language already would ask if they just forgot something. And then you want a set of things that a normal fluent adult would say as a matter of course, those little in between words, um's and oh's, most likely not in the dictionary.

Bootstrapping words


I don't understand.
I don't understand that word.
I don't know.
I don't think so.
I'm not sure.
I know.
I understand.

What? (I didn't hear you or I didn't understand)
Can you repeat that?
Can you say that slowly?
Can you say that again? Can you repeat that?
Can you explain that? Can you say that a different way?

What is that?
What do you call that?
What does X (a word in the new language) mean?
Is there a word for (some description in the new language)?
How do you say (some description in the new language)?
How do you say 'Y' (a word in your own language, not in the language being learned)? (if you are so lucky to already have someone bilingual as a teacher)
How do you spell 'X'? (if you are so lucky as to have a common writing system)

Of course.
Not at all, of course not.
Really?
Are you sure?
I can't tell the difference.
Do people really say that? Does everyone say that?

These should be a necessary part of the elementary language curriculum in any language. Also when translated into these other languages (the native language to be learned) they should be translated literally or formally, but have the corresponding way the natives say it.


For whatever language you're learning, find/translate/get these phrases in the new language and they'll speed up your fluency. The idea is to speak with awareness in the other language, to act as a native speaker would in their language. (OK, no native speaker would use many of these because of pragmatic concerns, not wanting to look like you don't understand).

Quasi-language




Quasi-language is utterances that are from the mouth but are not terribly dictionary oriented. To be a native speaker of a language, you should be able to 'say' these kinds of things (translated appropriately of course).

Uhhh, ummm - a 'spacer' between words while you think of what to say next instead of leaving a pregnant pause 
Uh-huh, yeah - very informal yes
Unh-unh, nuh-unh, mn-mm - very informal no
Ow - an exclamation of mild pain
Oops - if you dropped something or made a mistake
Tsk-Tsk - registering disapproval
Ah - registering understanding
Oh - registering mild surprise
Hey - to get someone's attention loudly
Shh - to tell someone to be quiet

Of course, some of these may be US/AmE particular and you 'just don't say it' in other 'languages'.

Saturday, October 31, 2015

A scientific taxonomy of ESP

This could just as well be framed as a taxonomy of magic or magical creatures or comic superpowers; you may disagree with details but the whole structure holds conceptually. There may be no actual facts involved (or maybe there are!), but the concepts are consistent. Also, I take this as a subset of the taxonomy of magic because there's (currently!) no scientific evidence but in the back of our heads we kind of feel like maybe we've experienced it or really really hope that there is some small ability there

First, let's define ESP (extrasensory perception) starting from examples, often being lucky enough that there are single English words that already capture the essence, and abstracting. There's clairvoyance (seeing the future), there's telepathy (perceiving someone's thoughts), mind-control (changing someone's thoughts by your own), speaking with the dead, telekinesis (moving objects with your mind), predicting random cards.

I'm setting an arbitrary boundary so that things we informally think are magical are not included (ghosts, gremlins, witches), that are 'obviously' unscientific and magic tricks (card tricks, optical illusions), which are intentionally supposed to seem magical but have a deterministic scientific explanation (astrology (depends supposedly directly on the location of the sun and planets)). These choices of mine are somewhat arbitrary. They could easily be included but then where do we stop (wait what about tarot and palm reading and tea leaves? what about entertainment magic, sleight of hand and actual tricks (ha ha that's hard to say right))

With these examples in mind, we can start to take apart what it means to be ESP and categorize all the kinds. The first thing to notice is that, along with perception, I am including action. So extrasensory perception or action is perceiving or doing things beyond our known senses. So we are well aware of seeing with our eyes and pushing with our hands; ESP is the ability to do those without currently known physiological organs. Presumably the organ will end up being the brain (the seat of thought), but maybe if we find out that we are able to see through the backs of cards using higher frequency receptors in our eyes (a deterministic scientific explanation) then this action will become a nonExtra Sensory Perception (NESP).

This brings up the tangent of making well formed categories. It is usually considered bad practice to have a subcategory, a sibling category, that is 'everything else that is not included'. For example, the category Vehicles could include Cars, Bikes, Planes, and NOS (Not Otherwise Specified). The latter category might cause difficulty because a sailboat will have to change category if a new subcategory of Vehicles, namely Boats, is created. (Note the difference between a category (eg Boats) and instances (sailboat), which of course could be generalized to become a category on its own)

A taxonomy of concepts forms a tree which expects all subtrees to be non-overlapping. Most collections of concepts end up having some overlaps, and this will be pointed out, but non-overlapping is a simplifying assumption that will make things easier to diagram.

- sensing
   - 'perceiving' events
      - clairvoyance, premonition - seeing events in the future, past, or remotely
         - guessing cards
         - predicting events
      - telepathy - knowing others' thoughts
         - mentalism - cold reading
         - channeling - communicating with spirits
            - seances - speaking with the dead (formerly actual people), knowing the thoughts of someone who has died
   - sensing auras - 'seeing' the personality of a person
   - out-of-body experience - astral projection

- acting
  - telekinesis - or psychokinesis, moving objects
     - levitation - raising objects
        - oneself - as in extreme yoga
        - somebody else
        - objects
     - making objects disappear
     - modifying objects
        - bending spoons
        - destroying and remaking things (watches, dollar bills)
     - pyrokinesis - starting fires (inspired/invented by fiction, Stephen King)
  - telepathy - transfer of thoughts, more than just sensing
    - sending thoughts, communicating
    - putting ideas in someone's head
    - mind control
    - body control

I've never defined magic or science, only working with them informally. The creation of the relations among these things helps us define our terms, putting things together that go together but avoiding conflicts and inconsistencies by separating differences.

This is an exercise in philosophy and taxonomy. That is, I'm just playing with words and our mental perception of them, mainly because science could be done on these things, and has, but it has just never panned out. So all I have to go on them is what we imagine. So this taxonomy is not (as currently known) about scientific things, but is itself scientific because people have ideas of what these individual concepts could mean and could disagree with the relations I have put among them. Note that I've really only put a subset relation (is-a) and extremely minimal comments.

The only practical argument against any of these abilities being real (or scientific) is that no one has used any of these things for anything other than those particular entertainments. That is, if ESP/magic were repeatable with other objects, we could use, for example, the spoon bending skill for other metals and substances in industrial manufacture. Or we could teach quadriplegics how to do small tasks requiring dexterity. Or communicate without telephones. Of course the counterargument which is not a counterargument is the ability of pickpockets to take personal objects without us knowing. Some 'magic' is possible, just not by the purported skills.

What's interesting about the above taxonomy is that most (serious) people don't really believe that any of these phenomena are real. This is counting angels on a pinhead, building castles in the sky. There is no there there. But we've drawn a perfectly coherent picture. And frankly, it could turn out that some of these are physically realizable, through some sort of deterministic, scientific process.



Friday, October 30, 2015

Language learners


What's happening with language learning:
  • One unknown word ruins a sentence: a single unknown word in a sentence can totally negate any meaning the sentence might otherwise give. If you have never heard a word before, for a native speaker you often have enough context and history to figure the part of speech, how it relates to the other words, who is doing what to whom. But to the non-native learner, it totally throws off everything. All the other words, which you previously know, may now have their meanings in question. And all the intellectual energy you're expending trying to figure out the unknown word is taken away from all the other words, making the known words, still shaky in this new language, even less sure. 
Statement by native speaker: "I went to the dumbledore to pick up some chicken and potato salad for the picnic"
Native speaker reaction: "dumbledore must be a grocery store."
Learner reaction: "Did you just call me a ... a bird?"
  • A learner is lenient, a native is strict: To a native speaker, there are lots of collections of words that are similar sounding and have similar meaning but are not the same. For example, all the various word forms of a conjugation "has, have, had" or cognates ", To the native speaker, these individual words are all very distinct. Using one instead of another is a glaring error, a discordant note, banging your thumb with a hammer obvious. To the language learner, they're kinda the same. To someone foreign to both, Italian and Spanish are a lot alike, you can sorta make half sense of both about the same. But of course to them they are mutually unintelligible (but can pick out a few words here and there). To the learner everything close is good enough. To the native the slightest hint of a difference is shockingly noticeable, strange, and almost unrecognizable. This works for all areas: pronunciation, syntax, word choice. 
"I want the grocery store"
Native speaker reaction: "Want? That makes no sense. Did you want something at the grocery store? Did you want a grocery store? I don't get it."
Learner reaction: "Did you get chicken and potato salad?"
    • Throw away step ladder: It seems universal in language teaching to start off with extremely simple sentences in the present indicative: "I read", "They eat". In English at least this is hardly ever used in practice. Surely there are short simple sentences that are actually used that can be taught.
    What is (might be) taught: "I went to the cinema Fridays."
    What normal people actually say: "I used to go to the movies every Friday."
        • Translationese: word for word translation is easy and often a sentence can preserve meaning from one language to the next with a constituent to constituent dictionary translation. But often "that's just not how they say it in X". Frankly in English that's just not they say it (see the throw away step ladder).
        What is taught: "That is correct"
        Natural: "Of course", "Right", "Yes", "Sure", "I guess so"
        • Style is not grammar: lots of rules are given (a consistent single rule to learn is much easier to remember than a more complex one) for which it is actually a style rule or a rule of register (formal vs informal). Also, most language teaching is academic and for a future business or academic use. In most languages (moreso but still a little in English) there is a big difference between the language called X in school and that called X at home. This can lead to a good language learner to be 'better' or excessively more formal than a native speaker. 
        What is taught: "I must go to the pharmacy momentarily to obtain some sundries"
        What people say: "I gotta go t'th'CVS 'n' pick up somethin' real quick"
        • Humor and sarcasm: anything other than the most literal will not be caught by the language learner. The learner probably has no idea that the same sounds can mean very different things depending on context (forgetting cultural background altogether). On the other hand, when a learner does notice a homophone, they will find it the most hilarious thing in the world but it will barely register with the native speaker.
        Lack of humor "Where do polar bears vote?"
        Native speaker: "The North Poll!"
        Learner: "I don't understand. Can polar bears vote?"

        Basic Pun "Where do polar bears vote?"
        Learner: "The North Poll! Ha ha! I get it! Because 'poll' sounds just like 'pole' but they're two different things, one is for ..."
        Native speaker: "Groan. Also, polar bears can't vote"


        The point is that someone learning a language is using all their mental energy to pick out the right sequence of words, to get the order right, pronunciation, to remember that one weird word, etc etc that it's the most unnatural thing in the word, and they sometimes miss the eventual meaning.

        Monday, October 26, 2015

        Vapnik says "Deep Learning is the Devil"...maybe

        Zach Lipton gave a summary of Vapnik's talk at Second Yandex School of Data Analysis conference (October 5-8, 2015, Berlin). Lipton wrote:
        Vapnik posited that ideas and intuitions come either from God or from the devil. The difference, he suggested is that God is clever, while the devil is not
        and

        Vapnik suggested that the devil appeared always in the form of brute force.
        and

        [Vapnik] suggested that the study of machine learning is like trying to build a Stradivarius, while engineering solutions for practical problems was more like being a violinist

        My interpretation of all this is that this is about the difference between science and engineering, or general vs specific. Coming up with a good general algorithm, I'm guessing Vapnik is thinking of SVMs or the idea of neural networks, is the study or science of ML, but most successes of Deep Learning (or really just particular and particularly large neural networks) come from the given design of the DL network.

        As to clever vs brute force, somehow the statement that can be extracted is that DL is not clever but devilishly brute force. I'm not sure how to make sense of this (I don't see how DL is more brute force that SVM or logistic regression or random forests). Unless all the work that must be done in engineering a good DL is in creating the topology of nodes; this is not automatic at all but needs a lot of cleverness to make a successful learner. But the DL part enables that cleverness (which would otherwise be impossible).

        Cleverness is not easily scalable; you can't just throw a whole bunch of extra nodes and arbitrary connections into a DL and hope it learns connections well, you have to  organize the layers well. Those details,, the needed to be clever is what slows down the scaling and I am guessing it what is 'devilish' about DL.

        This is all second hand and rewording of suggestions through someone's hearsay, and connecting dots that are barely mentioned and far apart. I'm totally putting words in his mouth, but this is what I expect Vapnik really means (or what I think Lipton thinks that Vapnik thinks, all telegraphically expressed). But really how much of anything is really not that?

        Friday, October 23, 2015

        Best Science Fiction Movies Ever

        My list of best science movies ever:
        • 2001: A Space Odyssey - the story is superior, the sets and effects still look modern.
        • Blade Runner  - again, the story is superior, the sets and effects still look modern. The attention to detail is amazing.
        • Star Wars - simple minded but so audacious. Too many false notes to be the top
        • The Matrix; Terminator - both stories excellent and very distinct but somehow similar. Hard to distinguish quality.
        • Jurassic Park - almost too commercial
        • Planet of the Apes - kitschy sets and costumes and dialog and acting (after the first one, they are all very amateurish), but rich in sci fi. This is the only one of mine that I think is very questionable. But I think it deserves a lot more credit.
        • Road Warrior (and Mad Max I) - 
        • Brazil; Memento (I know they're not considered scifi, but both belong here for me)
        • Too recent to judge reliably: Inception; Minority Report; Gravity; The Martian; District 9; Avatar; Children of Men; Source Code; Edge of Tomorrow; Looper; Predestination; Surrogates; Elysium; Limitless; Interstellar. I really liked all of these but I can't tell if they'll mean much to me later.
        • 12/16 Arrival

        I'm making little distinction between a single great movie in a series and the rest of the series. All of these refer mostly to the first ones in a series. I don't know what it is about Star Trek. The TV shows are way better than the movies; but the movies are somehow terrible (except as everyone agrees ST II: The Wrath of Khan). The modern J.J. Abrams reboots are enjoyable but nothing new and forgettable.

        A handful require a mention but just don't make the list: THX-138; A Boy and his Dog; The Andromeda Strain; Stalker; A Clockwork Orange; Sleeper; Galaxy Quest; were all important when I was younger, but not really anymore (well, every other line in Sleeper is memorable).

        I feel like I have to mention The Day the Earth Stood Still; Forbidden Planet; Metropolis only because they're near the top of everyone else's list. But I have to be honest and say they just feel so all around dated. 

        Most scifi movies are just junk: action-adventure schlock.

        I went through a few top 100 lists just to make sure I wasn't leaving any out. If the title is not in this list (and there's lots), sorry, I just didn't think enough of it. If you were to ask me about one not here, I'd probably say, "I suppose it was OK, but it just doesn't fit in my best ever list". For example Tron. I remember being very excited about that as a kid, but, despite its main idea, it's just not that great. Some people think Dune (by David Lynch) was epic, but I feel like they must have seen a different movie (I thought the film was terrible) and their opinion was colored too much by the book. Soylent Green, Westworld, and The Omega Man - these are of the same kitschiness as Planet of the Apes, but only have one note each (also Charlton Heston mostly; is that the problem?).

        The only one I feel bad about not putting on the list above is Wall-E. It's obviously well-done, there are a number of things in it that are memorable and prescient. But all I can say is "sorry animation".

        At the end all I can say is that this is not opinion, it is objective truth. So if you don't agree, either you're wrong, or I have made a transcription error.

        Motivated by the Skeptics Guide to the Galaxy Episode #536 (10/17/2015) their top 5 sci fi movies.

        Tuesday, October 20, 2015

        Cognitive computing is AI rebranded from the point of view of an app

        Artificial Intelligence is whatever it is about computing that is sort of magic. As users we don't know why exactly it works but it just does. As builders it's like the dumbest of magic tricks: fast hands, misdirection, brute force beforehand. Sure a lot of research has gone into clever math for it, but once you look behind the curtain, you realize the excitement of the builder is in fooling the user, passing the Turing test by whatever means necessary. (I exaggerate considerably for effect. There's lots of rocket science behind the curtain, but what distinguishes it from algorithms is that it embraces inexact heuristics rather than shunning them).

        AI doesn't have to be electronic. It could be mechanical, like a ball-bearing finite state machine that computes divisibility by three, or biological (like a chicken taught to do tictactoe). The chicken winning is magic. The brute force is determining the tictactoe decision tree then teaching the poor chicken. In the end, AI is mostly just computers.

        It is usually something that humans are only able to do: language, vision, and logic.

        AI is often used for just a heuristic such as the game 2048. The AI used to try to get a better score is essentially heuristics found by a good human player that were then coded as rules in a deterministic program. When you open the hood, there's no rocket science, it's just "always make a move that keeps the highest item in the corner", a human thought that is better than random and better than a beginner, but not perfect. An airline flight suggester is just (OK it is sorta rocket science) a special linear optimization problem (on a lot of data). Linear optimization is usually not considered AI but let's not quibble.

        AI usually means that some learning was involved at some point but usually that learning is not continuous, learning in operation. Probably some ML algorithm was run on a lot of past data to generate a rule and then set in stone (until another pass on more recent data updates the rule).

        The short history of AI is that it was invented/named in the late 50's, expected to solve all problems in a couple years in the mid 60's

        Cognitive computing is IBM's way of reintroducing AI to consumers. It's not AI, but it's not not AI. That is, it is AI dressed up as usable applications or modules that can be fit together to give the appearance that a human is behind it without having an actual human having to step in to do it. The usual list of properties that a cognitive computing app has are: awareness of context of the user, giving the user what they want before they ask for it, they learn from experience, deals well with ambiguity. But then it will probably also incorporate human language input or visual pattern recognition or thinking through a number of inference steps.

        Do you need an app that'll give you a new good tasting recipe for tacos? Deciding what's best is probably a good human task (but shhh we have an ML algorithm that figured out what are good ingredient combinations). Do you need an app that suggests to you good personalized travel plans? And now for something actually practical, do you need an app that will help discover cancer cures from buckets of EHR data?

        All of these are not the usual single narrow one-off AI apps (back up a trailer, distinguish cats from dogs in images, compete in rock paper scissor competitions).

        Without a doubt, cognitive computing is totally a hype/marketing term, new enough not to be some old over-used baggage-laden term like AI, not misleading (these are all sort of thinking apps) and vague enough to allow all sorts of companies to jump on the bandwagon with "Why yes, we've been doing cognitive computing before the term was invented!"

        But hype terms can be useful. Cognitive computing is a good label for engineering a combination of features, some that are traditional AI and some that are just good design that are becoming more obvious to have. Whether it's the machines doing the thinking in silicon, or the engineers doing some extra thinking in design, as long as the machine looks like it's reading your mind then that's a good app. It'll be useful if it catches on.

        Friday, October 16, 2015

        The Long Burning Hype of Hype Indicators

        (motivated by From Turing to Watson: The Long-Burning Hype of Machine Learning)

        Hype is bullshit. It is not true, but it is also not necessarily false. In fact it has only a tenuous connection to the true/false dichotomy/continuum. It is only exclamation.

        It is the hot-or-not score. It is the Time magazine weekly up or down cultural indicator. It is based on empty anecdotal perception, vaguely perceived frequency of mention or frequency of thought or coolness or I don't know what.

        It is barely a measure of anything other than the .

        The Gartner Hype Curve is also hype. It attempts to inform about the hype stage of many closely related items at once. But it turns out that is a piece of hype itself. The hype cycle is a well-hyped pseudo-scientific (non-evidenced based) proof-by-look-there's-a-picture.

        Here is the general pattern:




        It is very compelling. I have to be honest and say that that's exactly the timeline of how I think of things. At first I've just never heard of the thing. Then one mention.Then three in one day, then I hear and think of it all the time, then I get just sick of it, nauseated at the thought. Then it comes back as an accept everyday thing. Here's an example of a set of items from the 

        But... really? That's just a made up story. It seems to match what I think of as a story of popularity. It seems to match a good Hollywood drama: hero has early success and downfall and then third act of redemption.

        And it is just the vaguist notion of mood swings. And also what is the point besides entertainment, or schadenfreude or rooting for a comeback? 

        Look at the following. So sciency. Look at all the data points that you can follow year after year (the data viz is a bit hard to read the course of any particular item, but that's a minor quibble in comparison to the the central problems).

        The hype curve might be a useful thing, if only it measured something that is 1) coherent and 2) based on evidence. Introspection is a great inspiration but it is not measurable. 

        What is the meaning behind the graph? Also whatever the meaning what is the data underlying the graph?

        The easy answer is the source of the data. It is simply the 'educated' guess of Gartner analysts. Not an actual number. 

        Take any particular item. Does it follow the curve? Does it match the icon? Each label is a hype term, but has its own definitional problems. Each term can be vague, have multiple meanings, and have multiple incommensurate sources.

        Note also that the shape and timing of the graph is the same for all items. The different point icons are the only appeal to different scales for each item.

        So time, the x-axis, is incoherent (unmentioned context based for every item). Different thing might move along the supposed curve at different rates.

        But what about the y-axis? Is it popularity, that is, how often an item is mentioned (mentioned in tweets or on google)? or is it how successful an 'item' is (let's say quantitatively, money, earnings per year?) Even these ostensibly measurable concepts are problematic because of definition of terms (is one label the same or different than another).

        And once you nail down what the measurement should be, Gartner isn't doing any kind of such measurement, and there's no guarantee that the hype curve shape is a common pattern. It may be that an item's curve is up then down (then dead). Or it may have a steady rise. Or it may have multiple hype peaks at different scales. Or frankly it may have a curve like any stock, up down, steady, with most any pattern imaginable.

        If you want to make the hype curve useful. Pick a meaning (or meanings, heck go wild and have many types of hype) and then actually measure it. And only then will people... wel they won't accept that unconditionally, they'll also complain about the coherence of the concept measured and measuring difficulties. But at least it will be scientific evidenced based hype rather than just empty celebrity bullshit hype.

        Friday, October 9, 2015

        Kalman Filters = dynamic programming on linear systems for sensor accuracy

        Kalman filter is a method to increase the accuracy of a sensor in a linear system. (see this link for a visual explanation and derivation)

        The usual example is for the position of a space ship. You have the position/velocity of the ship and also an independent sensor of those. Both of those are somewhat iffy (usually assumed for continuous variables to be Gaussian). Using these two iffy things together, you can get a much more accurate approximation (smaller variance than both) of the current position/velocity.
        from bzarg


        The other ingredient of the method that makes it get called a Kalman filter is the the change in the sensed data is expected to be linear, so that all of this can be modeled using simple repeated matrix operations.

        Of course, this is not limited to dynamic mechanics but it makes the best presentation (because of the linear equations

        The point here (which is not to explain Kalman filters) is that the computational method of correction (abstracting away the matrices) is one of a one step recurrence relation (new sensor data at each step too) which is essentially dynamic programming and even better, you only need to know the most recent item.

        What's the point of a hold-out set?

        The purpose of a predictive model is to collect some sample data, calculate some function to help predict future unknown performance, hopefully with low error (or high accuracy).

        The classic statistical procedure takes the sample, a small subset of past data called the data or for later purposes the training set, does some rocket science on that set (say, linear regression), produces the model (some coefficients, some small machine that says yes or no or outputs a guess on a single new data point), and maybe also produces some extra measures that says how good or bad the fit is expected to be (correlation coefficient, F-test). And we're done. So many papers and studies have been done over the years that follow this pattern.

        But... what is a hold-out set? The modern way (not that modern) is to split the sample randomly into two parts, the training set (on which to do the classic part) and the test set or hold-out set to check. Run the model on all of the items in the test set and see how bad the fit. The test set is distinct from the training set because we want to validate on unseen data, we don't want to assume something we're trying to prove.

        Why do this? It seems like such a waste. Why in a sense throw away perfectly good sample data on a test when you could use it in making a more accurate model? Why in a sense test again when you can use that test data to train? More data is better, right?

        Well, you're not really throwing it away, but it does seem like a secondary, minor desire. After all, don't most statistical procedures compute some sort of quality measure on the entire set first? This desire not to 'waste' hard won sample data is very understandable; most of the labor in an experiment is not the statistics but in gathering the actual data.

        Of course one could weakly justify this test set by saying it gives more reliable quality statistics.

        The real desire for a holdout set is to combat overfitting. There are two sides to modeling: real life data is not perfect, the model is trying to get close to the rule behind the data, but it may go too far and get close to the data itself instead of the rule. The classic step gets us the first part, the modern step avoids going too far. A hint to the purpose is another name for the 'hold-out set which is validation set which name gives a better idea of its purpose. You create a model with the training set, and validate it with the validation set. You're validating your model, making sure that it does well what you claim does well. The first step in predictive modeling is to not underfit, to get close to reality that the data hopefully represents. The test or validation step is to make sure you don't overfit, get too close to the data at the expense of reality.

        So I've weakly justified the desire for some kind of hold-out/test set. But how does one actually choose this set? Obviously a random subset but what size? The primary issue is a balance between the model and the goodness of fit: with smaller training set, more variance in the model; with smaller test set, more variance in the stats. There's no hard and fast rule (80/20 is considered reasonable).There are a number of strategies to deal with this.


        • number not proportion - just make sure you have enough data points in each and after that proportion doesn't matter as much
        • resample- do the test a few times on random subsamples. This is the very general procedure bootstrap/jackknife
        • data partition and validate each as a test set against the rest - cross validation. This idea is to split the entire dataset into many samples and do the test/training on each set vs the rest. That is, all data is used as part of a training set and all as part of a test set at some point. There are many strategies here: leave one-out (LOOCV), where all but one is the training set and a single item is the test set, but do this for every single item in your data set. Under some models (like general linear regression models) you don't have to repeat the process n times because the math cancels out a lot (linearity is great!). Another method is k-fold CV where you split your data into k pieces (in practice often 5 or 10) and create a model on n-n/k items and validate on the, do that for each of these k pieces. It takes more time (k more times). LOOCV is essentially n-fold CV, so it is not efficient time wise when model creation takes a while (like for SVM)

        A lot of this ignores the issue of what to do if your validation set has bad performance. What is the statistically 'right thing to do' then? Do you rejigger things knowingly? How adaptive can you be and avoid p-hacking? I'll save that for later.