The joke is gonna be on all of you when I deploy a 256-core machine in production. :P
Then you'll modify it to be, "Have you heard of anyone running BrainFuck in production other than Nick P's BrainFuck-as-a-Service platform of questionable longevity?
Nice haha. I was just going to call it the BrainF interpreter with references to prototypes A-E. Ill say I put 256 "brains" in each FPGA or ASIC. I'll reference the performance of various machine learning algorithms. I'll have comparidons showing speedup over a Core 2 Duo I with less watts.
Fast-forward time some to see it becomes a post-mortum, legacy system, or acquired by Novell as their entry into AI. They promise it will be the success of Netware all over again. Audience at conference pauses at the ambiguity of that statement unsure if they should cheer or charge out the door.
Oh the funny part is I just slapped BrainFuck and a little creativity on top of the standard M.O. of the hardware accelerator industry. All the ones getting tons of VC money or revenue. Including DWave that showed the DWave-specific algorithm got a million times speed-up over implementing the DWave-specific algorithm on a general-purpose, barely-parallel CPU. Was that an innovation going a million times faster or a shoddy component going a million times slower? We don't know but they have reputable customers shelling out big cash. Lmao...
I wasn't aware of how flimsy the tech is. Guess I need to do more reading on the topic.
But it looks like there is a legitimate and interesting shift towards running stuff on FPGAs. Google, FB, and MS are all doing something with ML hardware acceleration.
I actually attended a talk a few weeks back by Microsoft's Doug Burger[1]. He has been leading a team that has created a low-latency FPGA network to accelerate stuff within MS. The eventual goal is to allow customers to take advantage of this distributed FPGA fabric to run custom firmware.
He said that FPGAs now run several of Bing's core search algorithms and Azure has some stuff running on FPGAs too. I forgot the exact performance gains, but it was somewhere around 2x for Bing with extremely stable response time even at insane server loads.
One interesting factoid is that they were able to translate Wikipedia in its entirety from English to Russian using 90% of the currently deployed FPGAs in around 100 ms. Insane stuff.
In second, this jumps out at me: "The first issue is that the problem instances where the comparison is being done are basically for the problem of simulating the D-Wave machine itself. There were $150 million dollars that went into designing this special-purpose hardware for this D-Wave machine and making it as fast as possible. So in some sense, it’s no surprise that this special-purpose hardware could get a constant-factor speedup over a classical computer for the problem of simulating itself."
Actually gives me an idea. Instead of comparison to BF competitors, I could actually just compare a massively-parallel, BF CPU to 256 interpreters communicating with each other through IPC running on a general-purpose computer. I'd show the CPU performed many times better. It's the closest thing I can think of to how D-Wave is doing benchmarking. The difference is $150 million is not in either my bank account or addition of transaction history.
"One interesting factoid is that they were able to translate Wikipedia in its entirety from English to Russian using 90% of the currently deployed FPGAs in around 100 ms. Insane stuff."
Didn't know about that project. Pretty cool. Yeah, the FPGA projects have been doing all kinds of stuff like that going back to at least the 90's from my reading. The speedups could be over fifty fold. Some claimed three digits. Other programs harder to parallelize & reduce... which is basically what they do on FPGA... might have under 100% speed up, tiny speed up, or even a loss if it was sequential algorithm vs ultra-optimized, sequential CPU like Intel's. The latest work, which started in 90's projects I believe, was to create software that automatically synthesizes FPGA logic from the fast path of applications in a high-level language then glues them into the regular application on a regular CPU. You can't get speed-up of actual hardware design but makes boosts easier if problem supports good synthesis. Tensilica is another example of a company whose Xtensa CPU is one that's customized... from CPU to compilation toolchain... to fit your application. Container people are compiling and delivering containers. Tensilica compiles and delivers apps with a custom CPU.
...Well, we can compile TIS-100 programs to BF, and then run it on your 256-core machine. Then the question will become, "Have you heard of anyone running BrainFuck in production, other than Nick P's BrainFuck-as-a-Service platform of questionable longevity, or Qwerty's TIS-100-as-a-service running atop the former?"
Interesting expansion that requires you to buy the new model that's 128-cores that each have an I/O MMU for added security of running the "corrupted" code without permanent damage. Includes actual ROM on-chip for trusted boot and accuracy of simulation. To keep customers happy, replacements for anyone with a support contract are sold at cost [with R&D included].
Yes, we would require a new model, as new IO primitves would be required. Might I suggest ^v/.! as the instructions for writing up, down, left, right and last IO ports, respectively? For reading, we could use &V\,*$ for up, down, left, right, any and last respectively.
We'll take whatever you suggest. So long as we have a DSL for writing our demo programs that buyers can understand and approve of before programmers see the actual primitives. ;)
Well, yes. With that, we can inplement most of the TIS-100 instruction set fairly easily. We may have to exclude JRO, the NIL psudoport, and restrict the use of labels, but it would be fairly complete. Better yet, these restrictions are totally undocumented, so programmers will have to either work with the primitives or cross their fingers that their code will compile each day.
"these restrictions are totally undocumented, so programmers will have to either work with the primitives or cross their fingers that their code will compile each day."
That's a great expansion it being undocumented. Worked wonders for Microsoft's strategy of lock-in. We're going to have to make sure the CPU and platform's API's are considered copywritten on top of that so they might not be... legally allowed... to clone or port it without paying us. We could collect royalties on a BrainFuck CPU for our lifetime plus an arbitrary number of years decided by Congress reps collecting bribes.
My pitch will target the enterprise DB, embedded, and military markets. I'll package it as something for highly-concurrent, real-time programming. Give them a language like ParaSail or Chapel that compiles to BrainFuck. Keep getting subcontracts to use it in criticsl, long-term projects. Eventually, the sham will come out but BrainFuck will be To Big To Fail (TM). Also leverage patent on key innovation forcing uptake of BrainFuck for at least 20 years.
I'm not sure it's worthy of a Bond villain but it's a start. ;)
Whatever, man. When all that falls on its face, CTOs will be googling for peeps with low level brainfuck debugging skills. But they can't hack BF themselves, so I'll ace the interview with no prep, and scoop a six fig salary. Remote, from Amsterdam. Part time, because, priorities.
Nah, they'll ask when they last had this problem. They'll learn that there were these migration tools developed to turn COBOL into C or Java with bug-for-bug compatibility. They'll pay some company like Semantic Designs to automate a conversion from BrainFuck to Rust since it will be popular on embedded systems by the time they're finished. The remoters will then spend all their time enhancing and debugging RustyBrainFuck while writing as much code in Rust as possible to hide the BrainFuck behind a neat interface. Building that was another enterprise project.
Unfortunately, you can't outrun software's hidden assumptions behind correct functioning forever. There will come a time, much like Y2K for COBOL, that their critical database will just loose everything if they don't make an internal change that requires understanding all the Rust, BrainFuck, Verilog, and analog components I used "because I was learning analog at the time."
I can't predict what they will do facing such a situation. I can tell you to start a brewery of fine Scotch that sends flyers to them around the time of the feasibility study. You'll make a killing. :)
It could be worse. No matter how bad the BF->Rust code that's generated is (it'd literally just be BF in a less compact syntax, most likely: it's pretty hard to convert idiom-wise), it's probably better than the Troff sources (as written by Joseph Ossanna).
Honestly, it'd probably be worth the effort to just write the Rust up front. The kind of technical debt you'd land in otherwise... euurgh.
That's been a tough question to answer recently. The greedy and the idiots are always on the top of my list. There's significant overlap in those in market segment that might buy this CPU. They strangely also have a large supply of cash that's been flowing for a long time. I wouldn't have guessed that based on what they taught me back in high-school about economics.
Then you'll modify it to be, "Have you heard of anyone running BrainFuck in production other than Nick P's BrainFuck-as-a-Service platform of questionable longevity?